Human Feedback Types
partialPairwise Preference
Directly usable for protocol triage.
"Large language models (LLMs) are currently applied to scientific paper evaluation by assigning an absolute score to each paper independently."
HFEPX · Eval paper review
Pujun Zheng, Jiacheng Yao, Jinquan Zheng, Chenyang Gu +5 more
Published
Mar 18, 2026
Citations
0
Trust level
Low
Usefulness score
40/100 (Low)
Extraction confidence
45% (Low)
Derived from extracted protocol signals and abstract evidence.
Rater population
Not reported
Signals refreshed
May 17, 2026
This paper is adjacent to HFEPX scope and is best used for background context, not as a primary protocol reference.
Use this as background context only. Do not make protocol decisions from this page alone.
Use this page for context, then validate protocol choices against stronger HFEPX references before implementation decisions.
Best use
Background context only
Use if you need
Background context only.
What to verify
Read the full paper before copying any benchmark, metric, or protocol choices.
Main weakness
The available metadata is too thin to trust this as a primary source.
Treat as adjacent context, not a core eval-method reference.
If you are doing eval pipeline work, start here
Large language models (LLMs) are currently applied to scientific paper evaluation by assigning an absolute score to each paper independently. However, since score scales vary across conferences, time periods, and evaluation criteria, models trained on absolute scores are prone to fitting narrow, context-specific rules rather than developing robust scholarly judgment. To overcome this limitation, we propose shifting paper evaluation from isolated scoring to collaborative ranking. In particular, we design a $\textbf{C}$omparison-$\textbf{N}$ative framework for $\textbf{P}$aper $\textbf{E}$valuation ($\textbf{CNPE}$), integrating comparison into both data construction and model learning. We first propose a graph-based similarity ranking algorithm to facilitate the sampling of more informative and discriminative paper pairs from a collection. We then enhance relative quality judgment through supervised fine-tuning and reinforcement learning with comparison-based rewards. At inference, the model performs pairwise comparisons over sampled paper pairs and aggregates these preference signals into a global relative quality ranking. Experimental results demonstrate that our framework achieves an average relative improvement of 21.8% over the strong baseline DeepReview-14B, while exhibiting robust generalization to five previously unseen datasets. Our code is available at https://github.com/ECNU-Text-Computing/ComparisonReview.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
Pairwise Preference
Directly usable for protocol triage.
"Large language models (LLMs) are currently applied to scientific paper evaluation by assigning an absolute score to each paper independently."
None explicit
Validate eval design from full paper text.
"Large language models (LLMs) are currently applied to scientific paper evaluation by assigning an absolute score to each paper independently."
Not reported
No explicit QC controls found.
"Large language models (LLMs) are currently applied to scientific paper evaluation by assigning an absolute score to each paper independently."
Not extracted
No benchmark anchors detected.
"Large language models (LLMs) are currently applied to scientific paper evaluation by assigning an absolute score to each paper independently."
Not extracted
No metric anchors detected.
"Large language models (LLMs) are currently applied to scientific paper evaluation by assigning an absolute score to each paper independently."
No benchmark or dataset names were extracted from the available abstract.
No metric terms were extracted from the available abstract.
Large language models (LLMs) are currently applied to scientific paper evaluation by assigning an absolute score to each paper independently.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
Detected: Pairwise Preference
Evaluation mode is explicit
No clear evaluation mode extracted.
Quality control reporting appears
No calibration/adjudication/IAA control explicitly detected.
Benchmark or dataset anchors are present
No benchmark/dataset anchor extracted from abstract.
Metric reporting is present
No metric terms extracted.