Skip to content
OpenTrain AIFor AI Companies

Researcher Tools

Human Feedback and Eval Paper Explorer

A focused feed for RLHF, preference data, rater protocols, agent evaluation, and LLM-as-judge research. Every paper includes structured metadata for quick triage.

Total papers: 501 Search mode: keyword Shortlist (0) RSS

Featured Papers

Popular high-signal papers with direct links to full protocol pages.

Weekly Eval Paper Digest

The top RLHF, evaluation, and human feedback papers — curated and summarized every Friday.

No spam. Unsubscribe anytime.

Start Here By Objective

Pick your immediate research objective and jump directly to high-signal pages, not generic search.

Scale Your Evaluation Team

Need human evaluators for your benchmark or preference study? OpenTrain sources pre-vetted domain experts into your annotation pipeline.

ImageEval 2026: Culturally Grounded Arabic Multimodal Evaluation

Samir Abdaljalil, Hunzalah Hassan Bhatti, Ahlam Bashiti, Farina Amir, Md Arid Hasan, Basel Mousi · Aug 31, 2026

Citations: 0

Match reason: Ranked by recency.

Score: 45% High protocol signal Freshness: Hot Status: Ready
Automatic Metrics General
  • We present an overview of the ImageEval 2026 shared task on culturally grounded Arabic multimodal evaluation.
  • We describe the task setup, datasets, evaluation procedure, and participating systems, and summarize the main results across the different tracks.
Open paper
Citations: 0

Match reason: Ranked by recency.

Score: 45% Moderate protocol signal Freshness: Hot Status: Ready
Pairwise Preference Automatic Metrics Coding
  • Healthcare workforce scheduling is an NP-hard optimization problem requiring simultaneous satisfaction of labor regulations, coverage requirements, employee preferences and cost objectives.
  • CP-SAT is evaluated on 18 instances: five synthetic hospital units (10-33 nurses), 10 INRC-II benchmarks (5-80 nurses, up to 8-week horizons) and 3 NRP-23 compatible instances (10-25 nurses) with cross-midnight Night shifts.
Open paper
Dense Clinical Contrasts Enhance Medical Knowledge Updating in Large Language Models

Yangmin Huang, Shu Quan, He Geng, Xin Ye, Qianyun Du, Zhiyang He · Aug 31, 2026

Citations: 0

Match reason: Ranked by recency.

Score: 45% Moderate protocol signal Freshness: Hot Status: Ready
Automatic Metrics Medicine
  • We introduce SEER-Bench, a temporally anchored oncology-staging benchmark curated from the latest versioned SEER Research Data release, and render identical medical update events from NCCN oncology guidelines into four supervision formats:…
Open paper
Citations: 0

Match reason: Ranked by recency.

Score: 42% Moderate protocol signal Freshness: Hot Status: Ready
Automatic Metrics General
  • We evaluate Hi-Q on three multi-hop QA benchmarks, primarily under full-corpus retrieval, where dependent evidence must be located among open-domain distractors rather than within a small annotated pool.
  • In this setting Hi-Q reaches 52.3 EM and 64.0 F1 averaged over the three benchmarks, ahead of the iterative retrieval baseline IRCoT by 15.1 EM / 18.2 F1 on that same average, and ahead of the graph-based RAG baseline PropRAG by 11.5 EM /…
Open paper
CHASE: How Content Ecosystems Are Reshaped When Ranking Is the Only Target

Qianwen Gao, Zichang Su, Yiwen Hou, Arlen Kumar, Leanid Palkhouski · Aug 31, 2026

Citations: 0

Match reason: Ranked by recency.

Score: 42% Moderate protocol signal Freshness: Hot Status: Ready
Simulation Env General
  • CHASE then iterates ranking, feature discrimination, rewriting, and evaluation over 20 rounds across different domains.
  • Quality-ranking alignment decreases in all six domains: from R0 to R20, the change in Spearman's rho ranges from -0.107 to -0.018, with a mean change of -0.068, which means documents closer to the ranking feature profile become less aligned…
Open paper
More Capable, Less Faithful: A Multilingual Analysis of Mathematical (Un)Solvability Detection in LLMs

Maria-Eleni Zoumpoulidi, Nikolaos Xiros, Georgios Paraskevopoulos · Aug 31, 2026

Citations: 0

Match reason: Ranked by recency.

Score: 42% Moderate protocol signal Freshness: Hot Status: Ready
Automatic Metrics MathMultilingual
  • To address this gap, we introduce the first multilingual benchmark of paired solvable and unsolvable mathematical problems, extending ReliableMath to French and Greek.
Open paper
Graph Evidence Is Not Enough: Diagnosing Native Decoder Use in Graph-Augmented LLMs

Xiaoyu Guo, Pengcheng Chen, Jiong Yu, Yi Lu, Yaohua Wang, Ziyang Li · Aug 31, 2026

Citations: 0

Match reason: Ranked by recency.

Score: 42% Moderate protocol signal Freshness: Hot Status: Ready
Automatic Metrics Medicine
  • Because the answer is a small integer and the target is purely topological, failure cannot be dismissed as open-ended generation or ambiguous evaluation.
Open paper
Whole-Slide Image Analysis under Realistic Few-Shot Annotation Protocols

Tiffanie Godelaine, Maxime Zanella, Karim El Khoury, Benoit Macq, Christophe De Vleeschouwer · Aug 31, 2026

Citations: 0

Match reason: Ranked by recency.

Score: 42% Moderate protocol signal Freshness: Hot Status: Ready
Automatic Metrics Medicine
  • Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Open paper
Learning Where Outcomes Change:Credit-Addressable Reasoning for Multimodal Geometry

Jiani Guo, Junjie Wang, Jie Wu, Pengxiang Zhao, Dongdong Zhang, Shaohan Huang · Aug 31, 2026

Citations: 0

Match reason: Ranked by recency.

Score: 45% Moderate protocol signal Freshness: Hot Status: Fallback
Automatic Metrics Long Horizon Coding
  • Across nine geometry benchmarks, CE-GRPO achieves an average accuracy of 76.04, outperforming Qwen3-VL-8B and trajectory-level GRPO by 8.09 and 3.43 points, respectively.
Open paper
Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions

Jaewoo Ahn, Junseo Kim, Hyunseo Kim, Heeseung Yun, Jaehyeon Son, Zsolt Kira · Aug 31, 2026

Citations: 0

Match reason: Ranked by recency.

Score: 45% Moderate protocol signal Freshness: Hot Status: Fallback
Llm As Judge Multi Agent General
  • Strategic deception by LLM and VLM agents has emerged as a central AI alignment and safety concern.
  • We introduce MineAmongUs, a 3D multimodal Among Us sandbox where imposter agents must deceive crewmates through joint verbal and non-verbal action.
Open paper
From Final Artifacts to Trajectories: Retrospective Process Supervision for Evidence-Grounded Long-Form Generation

Junjie Huang, Jiarui Qin, Di Yin, Weiwen Liu, Yong Yu, Xing Sun · Aug 31, 2026

Citations: 0

Match reason: Ranked by recency.

Score: 38% Sparse protocol signal Freshness: Hot Status: Ready
Long Horizon MathLaw
  • Trajectory data is getting more vital for training large language models for boosting the agentic abilities.
  • Experiments show that RetroGen improves grounding, faithful synthesis, and long-form evidence-seeking agent tasks.
Open paper
EvoSkill Injection: Red-Teaming Autonomous Skill Generation and Evolution in Self-Evolving Agents

Doyun Kim, Chanwoo Kim, Sugyeong Eo, Yeo-Chan Yoon, Chanjun Park · Aug 31, 2026

Citations: 0

Match reason: Ranked by recency.

Score: 42% Moderate protocol signal Freshness: Hot Status: Fallback
Red Team General
  • LLM-based agent systems increasingly adopt skill-based architectures to reduce repetitive reasoning costs and improve stable, efficient task execution.
  • Recent studies propose self-evolving agents that autonomously generate, refine, and reuse skills from past experiences to enable continuous capability evolution.
Open paper

Match reason: Ranked by recency.

Score: 35% Sparse protocol signal Freshness: Hot Status: Ready
General
  • Expanding the pretrained DFlash and DFlare drafters from block size 16 to 24 across Qwen3-8B and Qwen3-4B targets raises the per-prompt committed length on the high-ceiling benchmarks by a median of +0.8 tokens (up to +1.1).
  • The same expansion also lifts committed length on all seven benchmarks for Gemma-4-12B-IT, a different model family, by a median of +0.41 tokens (Arm A), and the full continuation-then-expand pipeline (Arm B) adds +0.29 to +0.98 tokens over…
Open paper

Match reason: Ranked by recency.

Score: 35% Sparse protocol signal Freshness: Hot Status: Ready
CodingMultilingual
  • Unlike most prior work that relies solely on English as the source language, we systematically assess different source-target language pairs and extend our evaluation to the more challenging, yet underexplored TASD task in cross-lingual…
Open paper
Towards Cognitive Process-Aware Proactive Writing Support

Masahiro Yoshida, Atsuya Kobayashi, Kei Tateno, Xiang 'Anthony' Chen · Aug 31, 2026

Citations: 0

Match reason: Ranked by recency.

Score: 35% Sparse protocol signal Freshness: Hot Status: Ready
General
  • Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Open paper