Human Feedback Types
partialPairwise Preference
Directly usable for protocol triage.
"Cold-start personalization requires inferring user preferences through interaction when no user-specific historical data is available."
HFEPX · Eval paper review
Avinandan Bose, Shuyue Stella Li, Faeze Brahman, Pang Wei Koh +5 more
Published
Feb 16, 2026
Citations
0
Trust level
Low
Usefulness score
40/100 (Low)
Extraction confidence
45% (Low)
Derived from extracted protocol signals and abstract evidence.
Rater population
Not reported
Signals refreshed
Feb 16, 2026
This paper is adjacent to HFEPX scope and is best used for background context, not as a primary protocol reference.
Use this as background context only. Do not make protocol decisions from this page alone.
Use this page for context, then validate protocol choices against stronger HFEPX references before implementation decisions.
Best use
Background context only
Use if you need
Background context only.
What to verify
Read the full paper before copying any benchmark, metric, or protocol choices.
Main weakness
The available metadata is too thin to trust this as a primary source.
Treat as adjacent context, not a core eval-method reference.
If you are doing eval pipeline work, start here
Cold-start personalization requires inferring user preferences through interaction when no user-specific historical data is available. The core challenge is a routing problem: each task admits dozens of preference dimensions, yet individual users care about only a few, and which ones matter depends on who is asking. With a limited question budget, asking without structure will miss the dimensions that matter. Reinforcement learning is the natural formulation, but in multi-turn settings its terminal reward fails to exploit the factored, per-criterion structure of preference data, and in practice learned policies collapse to static question sequences that ignore user responses. We propose decomposing cold-start elicitation into offline structure learning and online Bayesian inference. Pep (Preference Elicitation with Priors) learns a structured world model of preference correlations offline from complete profiles, then performs training-free Bayesian inference online to select informative questions and predict complete preference profiles, including dimensions never asked about. The framework is modular across downstream solvers and requires only simple belief models. Across medical, mathematical, social, and commonsense reasoning, Pep achieves 80.8% alignment between generated responses and users' stated preferences versus 68.5% for RL, with 3-5x fewer interactions. When two users give different answers to the same question, Pep changes its follow-up 39-62% of the time versus 0-28% for RL. It does so with ~10K parameters versus 8B for RL, showing that the bottleneck in cold-start elicitation is the capability to exploit the factored structure of preference data.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
Pairwise Preference
Directly usable for protocol triage.
"Cold-start personalization requires inferring user preferences through interaction when no user-specific historical data is available."
None explicit
Validate eval design from full paper text.
"Cold-start personalization requires inferring user preferences through interaction when no user-specific historical data is available."
Not reported
No explicit QC controls found.
"Cold-start personalization requires inferring user preferences through interaction when no user-specific historical data is available."
Not extracted
No benchmark anchors detected.
"Cold-start personalization requires inferring user preferences through interaction when no user-specific historical data is available."
Not extracted
No metric anchors detected.
"Cold-start personalization requires inferring user preferences through interaction when no user-specific historical data is available."
No benchmark or dataset names were extracted from the available abstract.
No metric terms were extracted from the available abstract.
Cold-start personalization requires inferring user preferences through interaction when no user-specific historical data is available.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
Detected: Pairwise Preference
Evaluation mode is explicit
No clear evaluation mode extracted.
Quality control reporting appears
No calibration/adjudication/IAA control explicitly detected.
Benchmark or dataset anchors are present
No benchmark/dataset anchor extracted from abstract.
Metric reporting is present
No metric terms extracted.