Skip to content
OpenTrain AIFor AI Companies

Researcher Tools

Human Feedback and Eval Paper Explorer

A focused feed for RLHF, preference data, rater protocols, agent evaluation, and LLM-as-judge research. Every paper includes structured metadata for quick triage.

Total papers: 501 Search mode: keyword Shortlist (0) RSS

Featured Papers

Popular high-signal papers with direct links to full protocol pages.

Weekly Eval Paper Digest

The top RLHF, evaluation, and human feedback papers — curated and summarized every Friday.

No spam. Unsubscribe anytime.

Start Here By Objective

Pick your immediate research objective and jump directly to high-signal pages, not generic search.

Scale Your Evaluation Team

Need human evaluators for your benchmark or preference study? OpenTrain sources pre-vetted domain experts into your annotation pipeline.

MuseCritic: Learning Multi-Aspect Song Rewards through Natural-Language Aesthetic Critiques

Jiabao Zhuang, Changhao Jiang, Hanchen Wang, Jiahao Chen, Zhixiong Yang, Zhenghao Xiang · Aug 12, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 58% High protocol signal Freshness: Warm Status: Ready
Pairwise PreferenceCritique Edit Automatic Metrics General
  • Long-form song generation models continue to improve in duration, structural integrity, and acoustic complexity, making reliable aesthetic rewards increasingly important for aligning these models with human preferences.
  • On the out-of-domain Music Arena benchmark with 733 preference pairs, it achieves the highest accuracy of 71.35%.
Open paper

Match reason: Matches selected tags (General).

Score: 58% High protocol signal Freshness: Warm Status: Ready
Demonstrations Automatic Metrics General
  • On the full GPQA Diamond benchmark (198 graduate-level science questions), majority voting reduces per-problem accuracy on a majority of problems for two instruction-tuned models from different families: 56.6% of problems for Qwen2.5-7B and…
Open paper
Automated grading of Linux/bash examinations using large language models: a four-level cognitive taxonomy approach

Manuel Alonso-Carracedo, Ruben Fernandez-Boullon, Pedro Celard, Francisco J. Rodriguez-Martinez, Lorena Otero-Cerdeira · Jul 2, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 58% Moderate protocol signal Freshness: Warm Status: Ready
Rubric Rating Automatic Metrics General
  • Gemini~3.0 Pro with rubric-guided prompting achieved the highest human-AI agreement (ICC(3,1) = 0.888, MAE = 0.10, Bland-Altman bias = -0.014).
  • These results show that question complexity is a reliable predictor of the difficulty LLMs face in grading accurately, and they establish a principled, taxonomy-based framework for determining which questions are suitable for AI-assisted…
Open paper
GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation

Ofir Ben Shoham, Shrutendra Harsola, Vignesh Subrahmaniam, Shravan Mohan, Yakov Gazman, Oded Vainas · Aug 12, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 55% Moderate protocol signal Freshness: Warm Status: Ready
Rubric Rating Llm As Judge General
  • Our reward is an LLM-as-a-judge rubric that scores each recommendation across multiple binary dimensions of advice quality, augmented with a safety gate for harm prevention.
  • Since LLM-based evaluation alone cannot confirm whether improvements reflect genuine business value rather than adaptation to the judge, we complement it with a judge-independent audit based on a standard doubly-robust Conditional Average…
Open paper
One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL

Simon Yu, Nicholas Tomlin, Marwa Abdulhai, Ximing Lu, Derek Chong, Abe Hou · Aug 12, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 58% Moderate protocol signal Freshness: Warm Status: Fallback
Simulation Env Multi Agent General
  • Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior.
  • Verbalized Sampling improves held-out success by up to 9% over single-simulator RL, and Co-Training pushes gains further to 14%; the human study shows similar gain on real users.
Open paper

Match reason: Matches selected tags (General).

Score: 58% Moderate protocol signal Freshness: Warm Status: Fallback
Llm As JudgeAutomatic Metrics General
  • Predictive-distribution entropy is a strong answer-selection rule in retrieval-augmented generation (RAG) for question answering: across five QA benchmarks, selecting the answer a frozen respondent LLM produces with the lowest answer-token…
  • LODESTAR uses reinforcement learning (GRPO) to train, once and offline, a polarizer -- a short fixed natural-language string inserted into the respondent's prompt and never into its weights, directing entropy so that entropy-based answer…
Open paper
ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents

Yutao Mou, Pengfei Yang, Zhe Yin, Zhangchi Xue, Xiaotian Luan, Dingyao Yu · Aug 12, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 58% Moderate protocol signal Freshness: Warm Status: Fallback
Simulation Env Tool Use General
  • Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states.
  • To bridge this gap, we propose **ToolHazard**, a scalable adversarial environment synthesis framework that reduces human engineering and supports expansion with additional seed domains and compute.
Open paper
Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs

Nimet Beyza Bozdag, Emre Can Acikgoz, Gokhan Tur, Dilek Hakkani-Tür · Aug 12, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 58% Moderate protocol signal Freshness: Warm Status: Fallback
Automatic Metrics Multi Agent General
  • As LLMs increasingly debate, advise, and think collaboratively with humans and each other, resistance to harmful persuasion becomes a core requirement for reliable behavior.
  • We formalize this threat as adversarial persuasion and introduce an adversarial reinforcement learning framework that trains persuader agents to change a target model's answer in a single interaction.
Open paper
Citations: 0

Match reason: Matches selected tags (General).

Score: 58% High protocol signal Freshness: Warm Status: Fallback
Automatic Metrics Long Horizon General
  • For LLM agents, however, the unit of observation is an interactive trajectory, where the model can ask clarifying questions, call tools, update state, and make intermediate decisions whose errors propagate to the final outcome.
Open paper
From Reasoning Depth to Reasoning Breadth: Evaluating Multi-Point Associative Reasoning in Large Language Models

Si'an Xie, Jiaxun Liu, Biao Yang, Wei Yuan, Fan Yang, Tingting Gao · Aug 11, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 58% High protocol signal Freshness: Warm Status: Fallback
Automatic Metrics Long Horizon General
  • We introduce MPAR-Bench, a bilingual English-Chinese benchmark that isolates reasoning breadth through multi-point associative reasoning.
  • We construct 1,000 items using a multi-agent clue-generation pipeline, embedding-based diversity filtering, and human verification.
Open paper
AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents

Xiangchen Cheng, Yunwei Jiang, Jianwen Sun, Zizhen Li, Chuanhao Li, Xiangcheng Cao · Jul 2, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 58% Moderate protocol signal Freshness: Warm Status: Fallback
Automatic Metrics Long Horizon General
  • Memory for a long-horizon LLM agent is a contract about what each future decision is allowed to see.
  • A public online benchmark of frontier LLMs on the same game reports zero wins at the lowest difficulty across five configurations, and the developer-reported human win rate at the same difficulty is 16%; the task is hard but not saturated.
Open paper
Citations: 0

Match reason: Matches selected tags (General).

Score: 55% Moderate protocol signal Freshness: Warm Status: Fallback
Pairwise Preference General
  • Through controlled experiments on two GPU families spanning the memory-bound to compute-bound regimes (A10G, ridge \approx 117 FLOP/byte; A100, ridge \approx 183 FLOP/byte), we demonstrate: (1) the inference-optimal vocabulary shifts 16x…
Open paper
Citations: 0

Match reason: Matches selected tags (General).

Score: 55% Moderate protocol signal Freshness: Warm Status: Fallback
Pairwise Preference General
  • Natural language user preferences provide an interpretable interface for LLM personalization.
  • To this end, we propose AlignXada, a training-free meta-learning framework that induces reusable textual refinement policies for adapting universal preference summaries to task-specific ones.
Open paper

Match reason: Matches selected tags (General).

Score: 55% Moderate protocol signal Freshness: Warm Status: Fallback
Critique Edit General
  • Open-ended revisions require a different measurement strategy because answer quality is graded, latent, and judged imperfectly.
  • We introduce an experimental protocol implemented across a pooled main peer-condition corpus and separately constructed decomposition corpora, allowing us to separate ordinary re-answering, candidate-content exposure, a bundled…
Open paper
EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments

Zhilin Wang, Han Song, Runzhe Zhan, Jusen Du, Jiacheng Chen, Tianle Li · Jul 2, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 55% Moderate protocol signal Freshness: Warm Status: Fallback
Simulation Env Long Horizon General
  • Autonomous agents are increasingly expected to improve executable policies through feedback, yet existing evaluations often collapse this process into a final score or confound it with open-ended software-engineering progress.
  • We introduce Autonomous Policy Evolution, a controlled evaluation setting in which a harness-model agent repeatedly edits an executable policy system under a fixed interaction budget.
Open paper

Match reason: Matches selected tags (General).

Score: 52% Sparse protocol signal Freshness: Warm Status: Fallback
Pairwise Preference General
  • Large language models are increasingly deployed in settings that require simultaneous adherence to multiple explicit constraints - reasoning structure, safety boundaries, output schemas.
  • We introduce Constraint Saturation Evaluation (CSE), a procedurally generated benchmark that systematically varies the number of simultaneous constraints (k), with every constraint scored by a deterministic, rule-based verifier and zero…
Open paper

Match reason: Matches selected tags (General).

Score: 52% Sparse protocol signal Freshness: Warm Status: Fallback
Pairwise Preference General
  • As users increasingly turn to Large Language Models (LLMs) for information and advice on political matters, particularly during election periods, the political preferences expressed by these systems have become a matter of public interest.
  • We demonstrate the framework through an Italian case study, providing a systematic analysis of LLM-generated political evaluations on italian parties and leaders.
Open paper
Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs

Vu Duc Anh, Nhat M. Hoang, Do Xuan Long, Cong-Duy Nguyen, Ponhvoan Srey, Luu Anh Tuan · Aug 12, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 52% Sparse protocol signal Freshness: Warm Status: Fallback
Pairwise Preference General
  • The first stage strengthens step-level reasoning via step-level preference optimization, while the second stage explicitly trains models to self-verify and self-correct.
  • Comprehensive in-domain and out-of-domain evaluations across multiple LLMs demonstrate that SFS-DPO and SFS-DPO-R consistently outperform prior step-level training baselines.
Open paper
Group Alignment-Induced Sycophancy: A Two-Sided Evaluation of Steerable Pluralistic Alignment

Haokai Zhao, Yunze Xiao, Weihao Xuan, Flora Salim, Benjamin Tag, Aditya Joshi · Aug 12, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 52% Sparse protocol signal Freshness: Warm Status: Fallback
Pairwise Preference General
  • Group alignment adapts a language model to a demographic group to produce responses that reflect the group's opinions, values, and preferences.
  • However, existing group alignment methods and evaluations focus only on how closely the model matches the group's opinions, overlooking the induced change in sycophantic behaviour.
Open paper