Skip to content
OpenTrain AIFor AI Companies

Researcher Tools

Human Feedback and Eval Paper Explorer

A focused feed for RLHF, preference data, rater protocols, agent evaluation, and LLM-as-judge research. Every paper includes structured metadata for quick triage.

Total papers: 340 Search mode: keyword Shortlist (0) RSS

Featured Papers

Popular high-signal papers with direct links to full protocol pages.

Weekly Eval Paper Digest

The top RLHF, evaluation, and human feedback papers — curated and summarized every Friday.

No spam. Unsubscribe anytime.

Start Here By Objective

Pick your immediate research objective and jump directly to high-signal pages, not generic search.

Scale Your Evaluation Team

Need human evaluators for your benchmark or preference study? OpenTrain sources pre-vetted domain experts into your annotation pipeline.

WebWorld: The Browser as a World Model for Self-Improving Web Code

Jiajun Wu, Jian Yang, Yaxin Du, Wei Zhang, Haowen Wang, Junhang Cheng · Aug 31, 2026

Citations: 0

Match reason: Matches selected tags (Coding).

Score: 65% High protocol signal Freshness: Hot Status: Ready
Critique Edit Simulation Env Web Browsing LawCoding
  • VLM-driven self-improvement of web code has a structural flaw: the model that proposes the repair is the model that judges it, and visual plausibility under that judge is a poor proxy for whether the page actually works.
Open paper
Citations: 0

Match reason: Matches selected tags (Coding).

Score: 65% Moderate protocol signal Freshness: Hot Status: Ready
Pairwise Preference Automatic Metrics Coding
  • Healthcare workforce scheduling is an NP-hard optimization problem requiring simultaneous satisfaction of labor regulations, coverage requirements, employee preferences and cost objectives.
  • CP-SAT is evaluated on 18 instances: five synthetic hospital units (10-33 nurses), 10 INRC-II benchmarks (5-80 nurses, up to 8-week horizons) and 3 NRP-23 compatible instances (10-25 nurses) with cross-midnight Night shifts.
Open paper

Match reason: Matches selected tags (Coding).

Score: 65% High protocol signal Freshness: Hot Status: Ready
Pairwise Preference Automatic Metrics Coding
  • We introduce OenoBench, a wine-domain knowledge benchmark of 3,266 multiple-choice questions across six pillars (regions, grape varieties, viticulture, winemaking, producers, business) and four difficulty tiers.
  • Evaluating sixteen frontier configurations, we find: (i) overall accuracy spans 53%-84%, led by o3 at 83.6%; (ii) reasoning-mode lift concentrates in DeepSeek R1 (+6.8pp) and is absent in Claude Opus and Gemini Pro; (iii) Anthropic shows…
Open paper
Stopping and Routing LLM Judge Panels

Bin Zhu, Yi Xie, Yanghui Rao · Aug 20, 2026

Citations: 0

Match reason: Matches selected tags (Coding).

Score: 65% High protocol signal Freshness: Hot Status: Ready
Pairwise Preference Llm As Judge MathCoding
  • LLM evaluation pipelines often have many candidate judges: general LLM-as-a-judge prompts, reward models, safety classifiers, confidence variants, and task-specific verifiers.
  • The deployment question is not only which judge is best, but which judges should be called, on which examples, and when panel construction should stop.
Open paper
Decomposing Wrong-Consensus Agreement in LLM Self-Consistency

Lizhuo Zhang, Mengmeng Tang, Chenfeng Long, Xiaoyong Tang, Xiang Luo · Aug 19, 2026

Citations: 0

Match reason: Matches selected tags (Coding).

Score: 65% High protocol signal Freshness: Hot Status: Ready
Pairwise Preference Automatic Metrics Coding
  • The mechanical reference is leak-free: each case's preference and accuracy are estimated from its other runs only.
  • A cross-system contrast at comparable aggregate accuracy contrasts near-complete mechanical agreement in the open-weights models against a larger preference-unexplained residual in the frontier family.
Open paper
Learning Where Outcomes Change:Credit-Addressable Reasoning for Multimodal Geometry

Jiani Guo, Junjie Wang, Jie Wu, Pengxiang Zhao, Dongdong Zhang, Shaohan Huang · Aug 31, 2026

Citations: 0

Match reason: Matches selected tags (Coding).

Score: 65% Moderate protocol signal Freshness: Hot Status: Fallback
Automatic Metrics Long Horizon Coding
  • Across nine geometry benchmarks, CE-GRPO achieves an average accuracy of 76.04, outperforming Qwen3-VL-8B and trajectory-level GRPO by 8.09 and 3.43 points, respectively.
Open paper
One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows

Zhuochun Li, Youngmin Ko, Ali Keramati, Nicola Ferri, Susana Palmaz Lopez Pelaez, Liang-Chun Tsai · Aug 20, 2026

Citations: 0

Match reason: Matches selected tags (Coding).

Score: 65% High protocol signal Freshness: Hot Status: Fallback
Automatic Metrics Tool Use Coding
  • Recent agent benchmarks increasingly ground evaluation in executable environments, from code repair to web navigation, app APIs, and function calling.
  • In this paper, we introduce Thinkingbox, a sandbox for tool-agent-user interaction that provides isolated MCP-compatible tool sessions, complete execution traces, and outcome evaluation over terminal backend state.
Open paper
Plans You Can Check: Verifier-Grounded Learning of an Open-Weight Planner for Executable Video-Editing

Haoyu Wang, Cheng Feng, Liuyang Bian, Ruiyang Huang, Lei Wei, Yafei Wen · Aug 26, 2026

Citations: 0

Match reason: Matches selected tags (Coding).

Score: 62% Moderate protocol signal Freshness: Hot Status: Fallback
Pairwise PreferenceRubric Rating Coding
  • A second stage, RefineCut-Evo, lets the student score its own repairs with the verifier and a task rubric and trains on high-margin preference pairs, so the final 8B planner runs in a closed verifier loop with no teacher calls at inference.
Open paper
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design

Yaxin Luo, Haobin Jiang, Jialv Zou, Xu Huang, Wenhao Yan, Haodong Li · Aug 13, 2026

Citations: 0

Match reason: Matches selected tags (Coding).

Score: 58% High protocol signal Freshness: Warm Status: Ready
Pairwise Preference Human Eval Long Horizon Coding
  • In this paper, we present AutoDesign, a framework that aligns with human design priors, where a meta-harness optimizer guides a code agent to recursively improve harness based on rollout feedback.
  • Across seven controlled code-agent-model configurations, integrating the learned DesignHarness consistently improves performance, increasing the average PosterBench Score from 54.99 to 67.39 (+12.4%).
Open paper
FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents

Yuhao Zhang, O. Ozan Koyluoglu, Thejas Venkatesh, Richard Diehl Martinez, Vishank Bhatia, Arash Alidoust · Aug 12, 2026

Citations: 0

Match reason: Matches selected tags (Coding).

Score: 58% Moderate protocol signal Freshness: Warm Status: Ready
Rubric Rating Llm As Judge Coding
  • We introduce FrontierFinance, a fully open benchmark of 220 expert-crafted queries and 11,543 source-attributed rubrics spanning six crucial use cases across the full investor workflow.
  • Evaluating frontier models and agent systems under a common harness restricted to publicly available data, we find that the tool harness, not the model alone, strongly shapes quality and efficiency; that Samaya's in-house system leads at…
Open paper
Principal Trait Analysis: Towards Deriving "Skills" in Human-AI Collaboration

Hunter McNichols, Kai Du, Andrew Lan · Aug 11, 2026

Citations: 0

Match reason: Matches selected tags (Coding).

Score: 58% Moderate protocol signal Freshness: Warm Status: Ready
Expert Verification Automatic Metrics Coding
  • Large Language Model-powered agents are increasingly used in the workplace via human-artificial intelligence (AI) collaboration.
  • We evaluate PTA on two human-AI collaborative coding datasets, an educational setting (students working with an AI tutor) and a professional setting (developers working with an AI coding agent).
Open paper

Match reason: Matches selected tags (Coding).

Score: 58% Moderate protocol signal Freshness: Warm Status: Ready
Pairwise Preference Automatic Metrics Coding
  • We evaluate SCDG on the PAN at CLEF benchmarks for generative plagiarism.
  • On a PAN 2025-derived pairwise benchmark, SCDG achieves 0.92 Precision, 0.97 Recall, and 0.94 F1, outperforming all baselines; on PAN 2026's multi-source retrieval task, it reaches 0.83 nDCG@10 and 0.96 Recall@100, surpassing all baselines.
Open paper
MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination

Saisha Shetty, Satvik Tripathi, Austin Lin, Colin Zhao, Theodore Kim, Don Enwerem · Aug 13, 2026

Citations: 0

Match reason: Matches selected tags (Coding).

Score: 55% Moderate protocol signal Freshness: Warm Status: Ready
Expert Verification Multi Agent MedicineCoding
  • We present Multi-Agent Reasoning and Coordination (MARC), an open-source framework that replaces monolithic LLM prompting with deterministic multi-agent orchestration for clinical reasoning.
  • MARC coordinates role-specialized agents for extraction, reasoning, answer generation, and evaluation, with explicit context passing and traceable intermediate outputs, enabling stage-wise failure attribution.
Open paper
Citations: 0

Match reason: Matches selected tags (Coding).

Score: 58% Moderate protocol signal Freshness: Warm Status: Fallback
Automatic Metrics Long Horizon Coding
  • Gist-based context compression---summarising older conversation history into compact representations---is a common approach in long-horizon language model agents, yet its effect on different types of memory retrieval is poorly understood.
  • The prompt modification recovers +0.314 [0.254, 0.375] judge accuracy on category-2 (temporal) questions in the matched set.
Open paper
Token Reduction Is Not Cost Reduction

Sarel Weinberger, Amir Hozez · Jul 13, 2026

Citations: 0

Match reason: Matches selected tags (Coding).

Score: 58% High protocol signal Freshness: Warm Status: Fallback
Automatic Metrics Long Horizon MedicineCoding
  • Token-reduction tools for coding agents are often evaluated by the number of tokens they remove, but token count alone does not determine end-to-end inference cost.
  • We evaluate three token-reduction approaches against an unmodified Claude Code baseline across controlled coding tasks, measuring provider-billed cost, task success, cache traffic, and agent behavior.
Open paper
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill

Zhuoyang Qian, Biao Wu, Yiran Wang, Chris D Yan, Desan Dai, Liangwei Zheng · Aug 12, 2026

Citations: 0

Match reason: Matches selected tags (Coding).

Score: 55% Moderate protocol signal Freshness: Warm Status: Fallback
Critique Edit Coding
  • We present Spark-to-Paper, an end-to-end research paper generation system implemented as thirteen composable skills inside an existing coding assistant, without requiring a separate agent platform or orchestration service.
Open paper
Index SLM Technical Report

Tianjiao Li, Lusheng Zhang, Shien He, Xiaojing Liu, Tianxing Yan, Mengran Yu · Jul 10, 2026

Citations: 0

Match reason: Matches selected tags (Coding).

Score: 52% Sparse protocol signal Freshness: Warm Status: Fallback
Pairwise Preference MathCoding
  • The series comprises four models: Index-1.9B-Base, a foundation model with 1.9 billion non-embedding parameters pre-trained on 2.8 trillion predominantly Chinese and English tokens; Index-1.9B-Pure, a control variant trained with an…
  • On a suite of standard benchmarks covering examination, reasoning, mathematics, and code, Index-1.9B-Base attains an average score of 64.92, competitive with or exceeding open models of several times its size.
Open paper