Skip to content
OpenTrain AIFor AI Companies

Researcher Tools

Human Feedback and Eval Paper Explorer

A focused feed for RLHF, preference data, rater protocols, agent evaluation, and LLM-as-judge research. Every paper includes structured metadata for quick triage.

Total papers: 501 Search mode: keyword Shortlist (0) RSS

Featured Papers

Popular high-signal papers with direct links to full protocol pages.

Weekly Eval Paper Digest

The top RLHF, evaluation, and human feedback papers — curated and summarized every Friday.

No spam. Unsubscribe anytime.

Start Here By Objective

Pick your immediate research objective and jump directly to high-signal pages, not generic search.

Scale Your Evaluation Team

Need human evaluators for your benchmark or preference study? OpenTrain sources pre-vetted domain experts into your annotation pipeline.

Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence

Ante Kapetanovic, Kemal Altwlkany, Andro Mercep, Tomislav Duricic, Emanuel Lacic · Aug 26, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 65% Moderate protocol signal Freshness: Hot Status: Ready
Critique Edit Llm As JudgeAutomatic Metrics General
  • Across 192,000 attempted evaluations (185,271 successful), seven out of the eight evaluated models have 95% task-stratified bootstrap intervals below zero for the total anchored-metadata effect on 20 fixed texts.
  • On categorical industry data with human-labeled ground truth, anchored metadata blocks 48% of error corrections and flips 10.18% of correct judgments toward an assigned wrong label, demonstrating the bias extends beyond numerical scoring to…
Open paper
Unfolding Scientific Papers into Multi-Turn Generation Trajectories for Continued Pre-Training

Qiankai Xu, Qiguang Chen, Zixin Su, Wenhao Huang, Yue Gao, Jiaheng Liu · Aug 26, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 65% Moderate protocol signal Freshness: Hot Status: Ready
Rubric Rating Long Horizon General
  • The same reverse construction extends to instruction data and evaluation.
  • Anchoring tasks in held-out papers yields PAW-Bench, an academic-writing benchmark whose tasks carry their own rubrics and checklists.
Open paper
Generative vs. Encoder Large Language Models for ASR Evaluation: A Comparative Study

Thibault Bañeras-Roux, Shashi Kumar, Driss Khalil, Sergio Burdisso, Petr Motlicek, Shiran Liu · Aug 26, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 65% Moderate protocol signal Freshness: Hot Status: Ready
Pairwise Preference Automatic Metrics General
  • While embedding-based metrics correlate better with human judgments, the respective roles of encoder and decoder-based Large Language Models (LLMs) remain underexplored.
  • This paper presents a comparative study of both families for ASR evaluation.
Open paper
Reflection Steering: Disentangling Reflection from Reasoning in Activation Space for Token-Efficient Inference

Jiarui Hu, Zhiyuan Wen, Xiaoyun Liu, Jiaxing Shen, Yu Yang · Aug 26, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 65% Moderate protocol signal Freshness: Hot Status: Ready
Critique Edit Automatic Metrics General
  • We conduct extensive experiments across two public benchmarks and three open-weight LLMs against state-of-the-art activation-steering baselines.
Open paper
ExecRubrics: Executable Tool-Augmented Rubrics for Verifiable and Efficient Long-Form Evaluation

Kaustubh D. Dhole, Charles L. A. Clarke, Eugene Y. Agichtein · Aug 23, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 65% High protocol signal Freshness: Hot Status: Ready
Pairwise PreferenceRubric Rating Automatic Metrics General
  • On three long-form response benchmarks -- HealthBench, HelpSteer, and ArgQuality -- we show that ExecRubrics can recover substantial preference signal without an LLM judge at evaluation time.
  • We show that incorporating external logic and resources from text processing libraries such as NLTK and spaCy can further improve preference accuracy.
Open paper
Move by Move: Measuring and Steering How LLMs Conduct Psychotherapy

Afonso Baldo, Hugo Pitorro, Areti Vassilopoulos, Anabela C. Areias, Maya D'Eon, Fabíola Costa · Aug 21, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 65% Moderate protocol signal Freshness: Hot Status: Ready
Expert Verification Automatic Metrics General
  • We introduce an ontology of ten therapeutic moves: compact, function-based categories grounded in the MULTI-60 inventory, validated through an annotation campaign with five licensed psychologists, and scaled with a judge-based approach that…
  • Applying it to real counseling transcripts and model-led sessions, we compare the move distributions between human clinicians and a panel of frontier models.
Open paper
Citations: 0

Match reason: Matches selected tags (General).

Score: 65% High protocol signal Freshness: Hot Status: Ready
Critique Edit Automatic Metrics Multi Agent General
  • Here, we introduce Tree-of-Concerns, a multi-agent framework that deploys specialized skeptic personas, each operating through a category-specific analytical lens, as parallel debate trees to extract unstated limitations from scientific…
  • Through experiments on ToC-Bench, our benchmark of 414 research papers with 1,905 unstated limitations, sourced from reviewer-reported weaknesses and follow-up citation critiques, we demonstrate that ToC improves precision by 79% and…
Open paper
LiLiCorr: Lightweight Likelihood Correlation of Parallel Drafts for Speculative Decoding

Matan Rusanovsky, Yoav Miron, Roy Uziel, Omer Belhasin, Ran Zilberstein, Maor Ashkenazi · Aug 20, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 65% Moderate protocol signal Freshness: Hot Status: Ready
Pairwise Preference Automatic Metrics General
  • Over the vanilla DFlash drafter, LiLiCorr raises acceptance length on every benchmark by 9 to 19%, while its scoring head accounts for about 2.8% of the per-block latency.
  • Against DFlash and two concurrent methods that also restore coherence at draft time, LiLiCorr delivers the highest throughput in 70 of 72 settings: nine benchmarks at two target sizes under greedy and temperature-one decoding, and a…
Open paper
Reward-Guided Autoregressive Graph Generation for Efficient Multi-Agent Communication Topology Design

Poomphob Suwannapichat, Boonyarit Changaival, Caesar Wu, Pascal Bouvry · Aug 20, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 65% Moderate protocol signal Freshness: Hot Status: Ready
Automatic Metrics Multi Agent General
  • LLM-based Multi-Agent Systems (MAS) achieve strong performance on complex reasoning tasks by coordinating multiple agents, but at the cost of substantial token consumption.
  • We address this limitation by introducing a Reward-Guided Autoregressive Graph Generation (RGA-Designer) inspired by Reinforcement Learning from Human Feedback (RLHF).
Open paper
Enhancing LLMs in Predictive Political QA with Semi-Structured Data

Yinan Liu, Zihan Zhou, Zichun Jin, Xinyu Wang, Bin Wang, Xiaochun Yang · Aug 21, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 62% Moderate protocol signal Freshness: Hot Status: Ready
Pairwise Preference Simulation Env General
  • We identify two complementary signals for predictive political QA: actor stances that capture issue-specific preferences, and high-order structure signals that capture indirect dependencies among political actors.
Open paper
When Text and Numbers Disagree: Evidence Arbitration in Large Language Models

Mattia Carletti, Edward Phillips, Fredrik K. Gustafsson, Patitapaban Palo, Lei Clifton, Danielle Belgrave · Aug 20, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 62% Moderate protocol signal Freshness: Hot Status: Ready
Pairwise Preference Tool Use General
  • To do so, we introduce a controlled synthetic benchmark in which latent risk trajectories generate both numerical time series and natural language summaries, allowing us to construct conflicts where exactly one evidence source is aligned…
  • Across open-weight instruction-tuned models, we find that arbitration behaviour is systematic rather than random: models exhibit distinct text-versus-number preferences, follow temporal recency more consistently than explicit reliability…
Open paper
ReliableRAG: Combating Misinformation in Retrieval-Augmented Generation via Reliability-Guided Reasoning Chains

Jinpu Jiang, Xuan Wu, Wenhao Song, Bo Yang, You Zhou, Hongwei Ge · Aug 26, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 65% Moderate protocol signal Freshness: Hot Status: Fallback
Automatic Metrics Long Horizon General
  • To address this limitation, we propose ReliableRAG, which, to the best of our knowledge, is the first reliability-driven framework that mitigates deceptive misinformation in multi-hop QA through fine-grained evaluation of individual…
Open paper
Multi-Agent Orchestration with the Common-Sense Reasoning Capabilities of LLMs for Autonomous Driving

Mehdi Azarafza, Faezeh Pasandideh, Ali Ehteshami Bejnordi, Stefan Henkler, Achim Rettberg · Aug 20, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 65% Moderate protocol signal Freshness: Hot Status: Fallback
Automatic Metrics Multi Agent General
  • While reinforcement learning and rule-based methods can provide effective control and safety mechanisms, their performance may degrade in situations requiring contextual reasoning.
  • The results demonstrate the potential of integrating LLM-based reasoning with conventional autonomous driving methods while retaining structured control and safety mechanism.
Open paper
MileGPO: Milestone Inference with Local Evidence for Graph-Based Policy Optimization of Long-Horizon LLM Agents

Bo Qian, Yuting Wu, Shuang Zeng, Huaiyu Wan, Dalin Zhang, Jiqiang Liu · Aug 20, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 65% High protocol signal Freshness: Hot Status: Fallback
Simulation Env Long Horizon General
  • Credit assignment is challenging in long-horizon agentic reinforcement learning, where supervision often comes only from final rewards.
Open paper

Match reason: Matches selected tags (General).

Score: 62% Moderate protocol signal Freshness: Hot Status: Fallback
Automatic MetricsSimulation Env General
  • To ensure that the summaries generated by SAraBERT achieve a high coverage of the document's main ideas, we propose Semantic Siamese Similarity, a novel evaluation metric that measures the level of similarity between two text inputs.
Open paper
Citations: 0

Match reason: Matches selected tags (General).

Score: 58% Sparse protocol signal Freshness: Hot Status: Fallback
Pairwise PreferenceRubric Rating General
  • In this survey, we introduce a Bayesian framework that defines constitutions as prior distributions P(R) over evaluation criteria and rubrics as conditional instantiations R_x \sim P(R|x).
  • Under this unified view, we present a taxonomy of rubric-guided RL along the prior-posterior axis, covering constitutional AI, instance-specific rubrics, process-level supervision, self-evolving rubrics, and their agentic and multimodal…
Open paper
Localize-Then-Decide Guarantees for LLM Judgments

Xinyu Li, Yi Zhou, Guanqun Cao, Zeyu Fu, Tianjin Huang, Gaojie Jin · Aug 26, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 58% Sparse protocol signal Freshness: Hot Status: Fallback
Pairwise Preference General
  • Large language models (LLMs) are increasingly used as evaluators to assess output quality and preference alignment, yet providing reliable guarantees of agreement with human judgments remains challenging.
  • Recent work introduces confidence-thresholding methods that provide such guarantees for pairwise comparisons, relying on the assumption that higher estimated confidence implies lower disagreement risk with humans.
Open paper
Personalized Privacy Control in LLMs via Attention Head Intervention

Junseok Kim, Nakyeong Yang, Kyomin Jung · Aug 21, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 58% Sparse protocol signal Freshness: Hot Status: Fallback
Pairwise Preference General
  • The rise of agentic AI enables LLMs to access diverse user data, raising critical privacy concerns.
  • To address this limitation, we introduce personalized privacy, which incorporates user-specific disclosure preferences into privacy control.
Open paper