Human Feedback Types
missingNone explicit
No explicit feedback protocol extracted.
"Generating multiple diverse and high-quality solutions is valuable for many applications, such as code-test generation and drug discovery."
HFEPX · Eval paper review
Yi Fang, Que Shen, Chengpeng Li, Boyi Deng +5 more
Published
Aug 31, 2026
Citations
0
Trust level
Low
Usefulness score
0/100 (Low)
Extraction confidence
20% (Low)
Derived from extracted protocol signals and abstract evidence.
Rater population
Not reported
Signals refreshed
Aug 31, 2026
This paper is adjacent to HFEPX scope and is best used for background context, not as a primary protocol reference.
Use this as background context only. Do not make protocol decisions from this page alone.
All signals on this page are inferred from the abstract only and may be inaccurate. Do not use this page as a primary protocol reference.
Best use
Background context only
Use if you need
Background context only.
What to verify
Validate the evaluation procedure and quality controls in the full paper before operational use.
Main weakness
This paper looks adjacent to evaluation work, but not like a strong protocol reference.
Treat as adjacent context, not a core eval-method reference.
If you are doing eval pipeline work, start here
Generating multiple diverse and high-quality solutions is valuable for many applications, such as code-test generation and drug discovery. However, Large Language Models (LLMs) tend to converge on a single high-confidence solution during inference, limiting exploration of alternative valid solution paths. Existing test-time methods promote diversity through tree-like search and prune semantically similar branches using response-level semantic embeddings. However, we find that such embeddings are easily confounded by linguistic and stylistic similarities, making it difficult to distinguish genuinely distinct solution paths. To address this, we introduce Answer Probing, which probes the potential answer an LLM would reach from an intermediate reasoning path. We demonstrate that the hidden states of probed answers more effectively differentiate distinct solution paths than semantic embeddings, and the perplexity of probed answers serves as a practical proxy for reasoning correctness. Based on these findings, we propose Answer Probing-Guided Tree Search (APTS), which guides the tree search by the probed answers' hidden state similarity and perplexity. Experiments on three reasoning tasks across two LLMs show that APTS consistently enhances solution diversity, demonstrating its effectiveness and robustness.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
None explicit
No explicit feedback protocol extracted.
"Generating multiple diverse and high-quality solutions is valuable for many applications, such as code-test generation and drug discovery."
None explicit
Validate eval design from full paper text.
"Generating multiple diverse and high-quality solutions is valuable for many applications, such as code-test generation and drug discovery."
Not reported
No explicit QC controls found.
"Generating multiple diverse and high-quality solutions is valuable for many applications, such as code-test generation and drug discovery."
Not extracted
No benchmark anchors detected.
"Generating multiple diverse and high-quality solutions is valuable for many applications, such as code-test generation and drug discovery."
Perplexity
Useful for evaluation criteria comparison.
"We demonstrate that the hidden states of probed answers more effectively differentiate distinct solution paths than semantic embeddings, and the perplexity of probed answers serves as a practical proxy for reasoning correctness."
No benchmark or dataset names were extracted from the available abstract.
Generating multiple diverse and high-quality solutions is valuable for many applications, such as code-test generation and drug discovery.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
No explicit human feedback protocol detected.
Evaluation mode is explicit
No clear evaluation mode extracted.
Quality control reporting appears
No calibration/adjudication/IAA control explicitly detected.
Benchmark or dataset anchors are present
No benchmark/dataset anchor extracted from abstract.
Metric reporting is present
Detected: perplexity