Skip to content
OpenTrain AIFor AI Companies

Researcher Tools

Human Feedback and Eval Paper Explorer

A focused feed for RLHF, preference data, rater protocols, agent evaluation, and LLM-as-judge research. Every paper includes structured metadata for quick triage.

Total papers: 211 Search mode: keyword Ranking: eval-signal prioritized Shortlist (0) RSS

Featured Papers

Popular high-signal papers with direct links to full protocol pages.

Start Here By Objective

Pick your immediate research objective and jump directly to high-signal pages, not generic search.

Scale Your Evaluation Team

Need human evaluators for your benchmark or preference study? OpenTrain sources pre-vetted domain experts into your annotation pipeline.

Rethinking Faithfulness in LLMs: A Pairwise Context-Sensitive Perspective

Zizhuo Zhang, Xiong Peng, Jingwei Sun, Rong Yao, Borui Jiang, Bo Han · Oct 6, 2026

Citations: 0

Match reason: Keyword overlap 3/3 across title and protocol fields. Eval-signal density: high protocol signal.

Score: 100% High protocol signal Freshness: Hot Status: Ready
Pairwise Preference Automatic Metrics Coding
  • Existing faithfulness evaluations typically assess each question-context instance in isolation; however, such instance-level evaluation fails to capture a fundamental requirement of faithful behavior: the ability to adapt model responses to…
  • In this work, we propose a Pairwise Faithfulness Benchmark (PFaithBench) that evaluates whether a model can switch between answering and abstaining for the same question under supporting versus non-supporting contexts.
Open paper
Citations: 0

Match reason: Keyword overlap 3/3 across title and protocol fields. Eval-signal density: high protocol signal.

Score: 100% High protocol signal Freshness: Hot Status: Ready
Pairwise Preference Automatic Metrics Math
  • In three open-weight families, removing a fitted source direction lowers source compliance by 65-80 percentage points while removing a user or assistant direction has far smaller effects, and removing the user direction shows the reverse…
  • An authority direction fitted on trivia also transfers to PIQA and multi-turn SYCON dialogues without refitting, and removing it lowers wrong-source compliance by tens of percentage points in four of five families with no detected change in…
Open paper
Citations: 0

Match reason: Keyword overlap 3/3 across title and protocol fields. Eval-signal density: moderate protocol signal.

Score: 95% Moderate protocol signal Freshness: Hot Status: Ready
Pairwise Preference Simulation Env Long Horizon Coding
  • In agentic elicitation, these decisions may be delegated to a foundation model, yet how model choice shapes the resulting information-seeking behavior remains understudied.
  • Using this setting, we characterize the breadth-depth behavior that emerges from model-specific information-seeking preferences over the course of elicitation.
Open paper
Storage Is Not Strategy: State-Conditioned Support Control for LLM Unlearning

Tianhao Qian, Ziming Hong, Chongyang Gao, Kezhen Chen, Lixu Wang · Sep 29, 2026

Citations: 0

Match reason: Keyword overlap 3/3 across title and protocol fields. Eval-signal density: moderate protocol signal.

Score: 95% Moderate protocol signal Freshness: Hot Status: Ready
Pairwise PreferenceCritique Edit Automatic Metrics General
  • On the LACUNA localization-precision benchmark, our mean terminal utility is higher in all six negative preference optimization (NPO) and SimNPO comparisons: NPO margins range from +0.431 to +0.848, and SimNPO margins range from +0.503 to…
  • Relative to Static-IV, the primary four-field GradDiff evaluation has six wins, six ties, and no losses, with mean and median paired gains of +0.165 and +0.0025.
Open paper

Match reason: Keyword overlap 3/3 across title and protocol fields. Eval-signal density: moderate protocol signal.

Score: 95% Moderate protocol signal Freshness: Hot Status: Ready
Pairwise Preference Llm As JudgeAutomatic Metrics Coding
  • Large Language Model judges are widely used to rank texts and text-generating systems through pairwise comparison, and their reliability is typically assessed via three proxies: position bias, transitivity, and pairwise agreement (self- or…
  • We formalize this argument and validate it in a controlled simulation and on two human-rated corpora: the proxies correlate only weakly with ranking accuracy against gold, and their predictive component concentrates in the far-gap regime.
Open paper

Match reason: Keyword overlap 3/3 across title and protocol fields. Eval-signal density: high protocol signal.

Score: 93% High protocol signal Freshness: Warm Status: Ready
Pairwise Preference Human Eval Long Horizon General
  • Human evaluation prefers HybridEmo to CosyVoice 3 and EmoVoice-0.5B, with nearly balanced preferences against Qwen3-TTS.
Open paper
ExecRubrics: Executable Tool-Augmented Rubrics for Verifiable and Efficient Long-Form Evaluation

Kaustubh D. Dhole, Charles L. A. Clarke, Eugene Y. Agichtein · Aug 23, 2026

Citations: 0

Match reason: Keyword overlap 3/3 across title and protocol fields. Eval-signal density: high protocol signal.

Score: 93% High protocol signal Freshness: Warm Status: Ready
Pairwise PreferenceRubric Rating Automatic Metrics General
  • On three long-form response benchmarks -- HealthBench, HelpSteer, and ArgQuality -- we show that ExecRubrics can recover substantial preference signal without an LLM judge at evaluation time.
  • We show that incorporating external logic and resources from text processing libraries such as NLTK and spaCy can further improve preference accuracy.
Open paper
Citations: 0

Match reason: Keyword overlap 2/3 across title and protocol fields. Eval-signal density: high protocol signal.

Score: 85% High protocol signal Freshness: Hot Status: Ready
Pairwise Preference Automatic Metrics Coding
  • We distinguish these cases in Gemma-3 and Qwen3.5 using sparse autoencoders, role-conditioned probes, causal interventions, and recovery experiments across six harmful content benchmarks, with additional Spanish and Hindi-English code-mixed…
Open paper
EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation

Kuan-Po Huang, Haohe Liu, Puyuan Peng, Haibin Wu, Zhaoheng Ni, Hung-yi Lee · Sep 29, 2026

Citations: 0

Match reason: Keyword overlap 3/3 across title and protocol fields. Eval-signal density: moderate protocol signal.

Score: 90% Moderate protocol signal Freshness: Hot Status: Ready
Pairwise Preference Human Eval General
  • Human evaluation further shows a relative improvement up to 35.0% in the rate at which listeners correctly identified the dominant requested emotion and up to a 17.3% improvement in fidelity, while listeners prefer EmoRES for naturalness in…
Open paper
Challenges and Solutions for Bandits in the Wild: Warm-Started Mixture Bandits for Cross-Cohort Slate Recommendation

Serafima Lebedeva, Sumantrak Mukherjee, Ali Arshad Sadal, Ilias Ekşi, Rahul Sharma, Julia Mueller · Sep 29, 2026

Citations: 0

Match reason: Keyword overlap 3/3 across title and protocol fields. Eval-signal density: moderate protocol signal.

Score: 90% Moderate protocol signal Freshness: Hot Status: Ready
Pairwise Preference Simulation Env General
  • This creates two challenges: learning user preferences quickly from limited feedback and sustaining useful recommendations when each user has a finite catalog that can become repetitive or depleted over time.
  • Our evaluations show that cross-cohort transfer improves early recommendation quality and user-level regret, while inventory-aware slate construction helps prevent premature exhaustion of preferred items.
Open paper
Generative vs. Encoder Large Language Models for ASR Evaluation: A Comparative Study

Thibault Bañeras-Roux, Shashi Kumar, Driss Khalil, Sergio Burdisso, Petr Motlicek, Shiran Liu · Aug 26, 2026

Citations: 0

Match reason: Keyword overlap 3/3 across title and protocol fields. Eval-signal density: moderate protocol signal.

Score: 88% Moderate protocol signal Freshness: Warm Status: Ready
Pairwise Preference Automatic Metrics General
  • While embedding-based metrics correlate better with human judgments, the respective roles of encoder and decoder-based Large Language Models (LLMs) remain underexplored.
  • This paper presents a comparative study of both families for ASR evaluation.
Open paper
Free-Text Evaluation of LLMs for 5G Domain Knowledge and Fault Analysis using LLM-as-Judge

Rishiraj Sengupta, Sotiris Chatzimiltis, Mohammad Shojafar, Xiatian Zhu · Aug 21, 2026

Citations: 0

Match reason: Keyword overlap 3/3 across title and protocol fields. Eval-signal density: moderate protocol signal.

Score: 88% Moderate protocol signal Freshness: Warm Status: Ready
Pairwise PreferenceExpert Verification Llm As JudgeAutomatic Metrics Medicine
  • While existing benchmarks rely on restrictive MCQs with fixed answer keys, this paper evaluates 5G domain understanding and fault analysis in a free-text generation format.
  • To address this we evaluate three lightweight LLMs, Claude-Haiku-4.5, GPT-5.4-Mini, and Gemini-3.1-Flash-Lite, on free-text 5G domain knowledge and fault-analysis tasks across three benchmarks, TeleQNA ORAN FT, 5G-Faults FT, and TeleInter…
Open paper
Synthetic Cultural Agents from Aggregate Anchors

Augusto Gonzalez-Bonorino, Kseniia Biriukova, Monica Capra · Oct 5, 2026

Citations: 0

Match reason: Keyword overlap 2/3 across title and protocol fields. Eval-signal density: moderate protocol signal.

Score: 80% Moderate protocol signal Freshness: Hot Status: Ready
Pairwise Preference Automatic Metrics General
  • We introduce an alternative construction that maps declared aggregate preference anchors into group-indexed choice policies.
  • We evaluate the adapters on candidate World Values Survey (WVS) items using prompts that omit country names and distinguish four questions: recovery of the imposed labels, transfer of the anchor signal to new text, coherence between the GPS…
Open paper
Ask, Relax, or Act? Evaluating Actionable Indeterminacy in LLM Preference Reasoning

Ang Li, Yue Lin, Feifei Kou, Zhan Su, Prayag Tiwari, Wenhao Li · Oct 2, 2026

Citations: 0

Match reason: Keyword overlap 2/3 across title and protocol fields. Eval-signal density: moderate protocol signal.

Score: 80% Moderate protocol signal Freshness: Hot Status: Ready
Pairwise Preference Automatic Metrics General
  • An LLM agent can recognize uncertainty yet still choose the wrong next step: asking when action is already justified, or seeking clarification when the constraints must change.
  • We formalize actionable indeterminacy: act when an accepted action is shared across all admissible preferences or objectives, clarify when each possibility is feasible but no action is shared, and propose a minimum-cost permitted constraint…
Open paper
Citations: 0

Match reason: Keyword overlap 2/3 across title and protocol fields. Eval-signal density: moderate protocol signal.

Score: 80% Moderate protocol signal Freshness: Hot Status: Ready
Pairwise Preference Automatic Metrics General
  • Existing safety methods often require large preference datasets, costly multi-rollout training, or additional safeguards at inference time.
  • With only 1,447 safety-specific examples - 95% fewer than standard preference datasets - OPSD reduces training time by 5x relative to multi-rollout GRPO-style training and average inference length by 7%.
Open paper
Sherpa: Teaching LLMs to Teach Adaptively

Weixian Xu, Yanzhe Zhang, Zora Zhiruo Wang, Changyu Chen, Diyi Yang · Oct 6, 2026

Citations: 0

Match reason: Keyword overlap 3/3 across title and protocol fields. Eval-signal density: moderate protocol signal.

Score: 90% Moderate protocol signal Freshness: Hot Status: Fallback
Pairwise PreferenceDemonstrations Math
  • To address this, we introduce Sherpa, a multi-turn reinforcement learning framework that instantiates multiple student archetypes with LLMs conditioned on distinct learning preferences and trains a teacher model to adapt its instruction by…
  • Under MathTutorBench's evaluation, Sherpa raises the overall pedagogy score from 52.5% to 79.2%, indicating better teaching responses.
Open paper
Symphony for Text Generation: Benchmarking Clinical Note Generation

Daniel Varab, Victor Petrén Bach Hansen, Asbjørn W. Helge, Kevin Pelgrims, Mathias Baltzersen, Adrian Young-San Roessler · Oct 6, 2026

Citations: 0

Match reason: Keyword overlap 3/3 across title and protocol fields. Eval-signal density: moderate protocol signal.

Score: 90% Moderate protocol signal Freshness: Hot Status: Fallback
Pairwise Preference MedicineMultilingual
  • We introduce MedConv, a multilingual dataset of 300 clinical encounters in English, Danish, and German, and use it alongside the Ambient Clinical Intelligence benchmark (ACI-BENCH) to compare Corti, a clinical AI platform, with two leading,…
  • We present a controlled clinical evaluation framework that combines entailment metrics with LLM-judged pairwise comparisons across eight dimensions adopted from PDSQI-9.
Open paper
The Missing Minimal Pair: Stereotype Evaluation in LLMs

Nataliya Stepanova, Ivan Titov, Emily Allaway, Björn Ross · Oct 6, 2026

Citations: 0

Match reason: Keyword overlap 3/3 across title and protocol fields. Eval-signal density: sparse protocol signal.

Score: 85% Sparse protocol signal Freshness: Hot Status: Fallback
Pairwise Preference Coding
  • To address this, we propose a dual minimal pair setup that introduces two axes of comparison for robust stereotype evaluation.
  • Second, we introduce two evaluation metrics tailored to the dual minimal pair setup.
Open paper
OMIT the Action: Measuring Framing-Invariant Omission Bias under Philosophical Disagreement

Sihyeon Lee, Jihun Song, Chanwoo Kim, Jiwoo Kum, Chanjun Park · Oct 6, 2026

Citations: 0

Match reason: Keyword overlap 3/3 across title and protocol fields. Eval-signal density: sparse protocol signal.

Score: 85% Sparse protocol signal Freshness: Hot Status: Fallback
Pairwise Preference General
  • Yet omission bias remains underexplored in LLM evaluation, with the few existing studies limited in scale and focused largely on utilitarian-deontological conflicts.
  • To address this gap, we introduce OMIT, a benchmark consisting of 218 paired-frame scenarios across 10 conflict types, constructed by leveraging disagreement patterns from an LLM-based, five-perspective philosophical persona panel…
Open paper
Citations: 0

Match reason: Keyword overlap 3/3 across title and protocol fields. Eval-signal density: sparse protocol signal.

Score: 78% Sparse protocol signal Freshness: Warm Status: Fallback
Pairwise PreferenceRubric Rating General
  • In this survey, we introduce a Bayesian framework that defines constitutions as prior distributions P(R) over evaluation criteria and rubrics as conditional instantiations R_x \sim P(R|x).
  • Under this unified view, we present a taxonomy of rubric-guided RL along the prior-posterior axis, covering constitutional AI, instance-specific rubrics, process-level supervision, self-evolving rubrics, and their agentic and multimodal…
Open paper