Skip to content
OpenTrain AIFor AI Companies

Researcher Tools

Human Feedback and Eval Paper Explorer

A focused feed for RLHF, preference data, rater protocols, agent evaluation, and LLM-as-judge research. Every paper includes structured metadata for quick triage.

Total papers: 501 Search mode: keyword Shortlist (0) RSS

Featured Papers

Popular high-signal papers with direct links to full protocol pages.

Weekly Eval Paper Digest

The top RLHF, evaluation, and human feedback papers — curated and summarized every Friday.

No spam. Unsubscribe anytime.

Start Here By Objective

Pick your immediate research objective and jump directly to high-signal pages, not generic search.

Scale Your Evaluation Team

Need human evaluators for your benchmark or preference study? OpenTrain sources pre-vetted domain experts into your annotation pipeline.

PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization

Boryeong Cho, Sumyeong Ahn, Se-Young Yun · Aug 31, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 65% Moderate protocol signal Freshness: Hot Status: Ready
Pairwise Preference Automatic Metrics General
  • To address this, we propose Posterior Label Correction DPO (PLC-DPO) to robustly optimize preferences by routing each pair's training signal as a clean, flip, or tie case.
  • Across 57 dataset-model-benchmark cells, PLC-DPO obtains the best mean win rate against DPO (60.5 vs.
Open paper
ScienceArena: Benchmarking LLMs on Latest Scientific Olympiad Competitions

Guangxiang Zhao, Qilong Shi, Xusen Xiao, Wenpu Liu, Yaoming Li, Linfeng Hao · Aug 31, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 65% High protocol signal Freshness: Hot Status: Ready
Rubric Rating Llm As Judge Long Horizon General
  • Benchmark saturation and data contamination increasingly obscure genuine scientific reasoning in frontier LLMs.
  • We introduce ScienceArena, an olympiad-style benchmark from thirteen public science competitions in physics, chemistry, and biology, including IPhO and IChO 2025--2026, IBO 2023, USAPhO 2026, and USNCO 2025.
Open paper
PAVE: Predictive Alignment and Value-Guided Evolution for World-Action Policies

Botong Zhao, Fang Yu, Tim, Senhua Zhu, Xinyuan Chen, Yue Lu · Aug 31, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 65% Moderate protocol signal Freshness: Hot Status: Ready
Demonstrations Simulation Env Long Horizon General
  • Across the three simulation benchmarks, \method achieves the strongest overall performance while preserving the direct actor's online execution path.
Open paper
Citations: 0

Match reason: Matches selected tags (General).

Score: 65% High protocol signal Freshness: Hot Status: Ready
Rubric Rating Llm As Judge Multi Agent General
  • Multi-Agent Debate (MAD) has been widely adopted to improve LLM-based evaluation by prompting multiple agents to negotiate and reach a consensus.
  • However, for subjective rubric-based scoring, inter-agent agreement does not guarantee alignment with human judgments.
Open paper

Match reason: Matches selected tags (General).

Score: 65% Moderate protocol signal Freshness: Hot Status: Ready
Pairwise Preference Automatic Metrics General
  • Centering better resolves subset diversity but loses this useful token-wise preference, revealing that diversity and distinctiveness are entangled in the raw geometry.
  • Based on this analysis, we propose the Centered Geometry Pruner (Cen-Prune), which measures subset diversity using centered cosine similarity while retaining raw-space distinctiveness as a complementary token-wise preference.
Open paper
PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback

Xueqing Wu, Ashwin Balasubramanian, Bingxuan Li, Dawei Zhu, Kai-Wei Chang, Yale Song · Aug 31, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 65% High protocol signal Freshness: Hot Status: Ready
Pairwise PreferenceCritique Edit Simulation Env Multi Agent General
  • To bridge this gap, we present MTPaperBananaBench, a benchmark for multi-turn diagram generation containing 292 images annotated with 3,518 user requirements.
  • To address these issues, we introduce PaperBanana-Interact, a multi-agent system that refines diagrams via an internal critique-and-refine loop.
Open paper
SimCRAFT: Distilling Remote Sensing Agents via Synthetic Trajectories and Contextual Retrieval-Augmented Fine-Tuning

Haoran Wang, Jing Yao, Xu Yang, Zeqing Wang, Yang Zhang, Pedram Ghamisi · Aug 31, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 62% Moderate protocol signal Freshness: Hot Status: Ready
Expert Verification Long Horizon General
  • The unprecedented surge in Earth observation data volume and diversity has exposed a critical bottleneck for traditional manual workflows, catalyzing the emergence of Remote Sensing (RS) Agents.
  • However, the practical deployment of these advanced agents is severely hindered by their heavy reliance on large-scale general-purpose LLMs, which lack deep domain expertise and impose prohibitive infrastructure demands.
Open paper
SwarmBench: Can Large Language Models Act as Agent Swarm Orchestrators?

Jinshan Gao, Zhuoran Jin, Tianyi Men, Kang Liu, Jun Zhao · Aug 31, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 65% High protocol signal Freshness: Hot Status: Fallback
Automatic Metrics Multi Agent General
  • Large language model-based multi-agent systems are evolving from fixed interaction topologies toward dynamically orchestrated Agent Swarms.
  • We propose SwarmBench, a benchmark that evaluates model performance from multiple perspectives, including accuracy, efficiency, cost, and process quality.
Open paper
Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions

Jaewoo Ahn, Junseo Kim, Hyunseo Kim, Heeseung Yun, Jaehyeon Son, Zsolt Kira · Aug 31, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 65% Moderate protocol signal Freshness: Hot Status: Fallback
Llm As Judge Multi Agent General
  • Strategic deception by LLM and VLM agents has emerged as a central AI alignment and safety concern.
  • We introduce MineAmongUs, a 3D multimodal Among Us sandbox where imposter agents must deceive crewmates through joint verbal and non-verbal action.
Open paper
DASC: Decay-Aware State Compression for Hybrid Linear-Attention Serving

Yanqi Yu, Pingwei Sun, Jianchao Tan, Tao Zhang, Yuchen Xie, Xunliang Cai · Aug 31, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 65% Moderate protocol signal Freshness: Hot Status: Fallback
Automatic Metrics Long Horizon General
  • Across retrieval and end-to-end reasoning benchmarks on Kimi-Linear, conservative DASC configurations remain close to full caching while compressing KDA recurrent state checkpoints by 2.63\times.
Open paper
Trajectory-Initialized Neural Double Q-Routing for Large-Scale Overhead Hoist Transport Systems

Cheng Gu, Qiusheng Zhao, Anbang Liu, Shaochong Lin, Max Z. J. Shen · Aug 31, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 62% Moderate protocol signal Freshness: Hot Status: Fallback
Simulation Env Long Horizon General
  • Large-scale industrial robot fleets share constrained physical infrastructure, making vehicle travel times dependent on safety separation, intersection access, downstream blocking, and station contention.
Open paper
EvoSkill Injection: Red-Teaming Autonomous Skill Generation and Evolution in Self-Evolving Agents

Doyun Kim, Chanwoo Kim, Sugyeong Eo, Yeo-Chan Yoon, Chanjun Park · Aug 31, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 62% Moderate protocol signal Freshness: Hot Status: Fallback
Red Team General
  • LLM-based agent systems increasingly adopt skill-based architectures to reduce repetitive reasoning costs and improve stable, efficient task execution.
  • Recent studies propose self-evolving agents that autonomously generate, refine, and reuse skills from past experiences to enable continuous capability evolution.
Open paper
Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents

Hanlin Tian, Minhao Li, Yu Mi, Sihan Zhu, Zhao Yang, Yuxiang Wang · Aug 31, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 62% Moderate protocol signal Freshness: Hot Status: Fallback
Rubric Rating General
  • Professional agent tasks often depend on conventions that are absent from public corpora, yet benchmarks rarely control whether an agent has access to those conventions.
  • Across fifteen calibration tasks, one frontier agent configuration achieves a 68.0% pass rate with the artefact and 0% without it; on one task, a plausible but incorrect artefact also yields 0% across five trials.
Open paper
Co-Evolving Actor-Conditioned Critics for Non-Verifiable Generation

Jinyoung Kim, Muhammad Khalifa, Lajanugen Logeswaran, Jaekyeom Kim, Moontae Lee, Honglak Lee · Aug 31, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 58% Sparse protocol signal Freshness: Hot Status: Fallback
Pairwise PreferenceCritique Edit General
  • We use this reward to train an actor-tailored critic with GRPO, and use critique-guided refinements to construct DPO preference pairs for the actor, forming a co-evolving critic-actor loop where the critic adapts to the actor's changing…
Open paper
When Errors Become Memories: Causal Pathway Tracing in Multi-Turn Memory-Augmented LLMs

Shuyao Xiao, Shengling Wang, Xuan Chen, Ke Chao, Ming Cui, Feifei Qian · Aug 31, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 58% Sparse protocol signal Freshness: Hot Status: Fallback
Pairwise Preference General
  • Error influence is evaluated at four levels: memory retention, natural responses, targeted diagnostic probing, and probability-level error preference.
Open paper