Skip to content
OpenTrain AIFor AI Companies

Researcher Tools

Human Feedback and Eval Paper Explorer

A focused feed for RLHF, preference data, rater protocols, agent evaluation, and LLM-as-judge research. Every paper includes structured metadata for quick triage.

Total papers: 125 Search mode: keyword Shortlist (0) RSS

Featured Papers

Popular high-signal papers with direct links to full protocol pages.

Weekly Eval Paper Digest

The top RLHF, evaluation, and human feedback papers — curated and summarized every Friday.

No spam. Unsubscribe anytime.

Start Here By Objective

Pick your immediate research objective and jump directly to high-signal pages, not generic search.

Scale Your Evaluation Team

Need human evaluators for your benchmark or preference study? OpenTrain sources pre-vetted domain experts into your annotation pipeline.

Stopping and Routing LLM Judge Panels

Bin Zhu, Yi Xie, Yanghui Rao · Aug 20, 2026

Citations: 0

Match reason: Matches selected tags (Math).

Score: 65% High protocol signal Freshness: Hot Status: Ready
Pairwise Preference Llm As Judge MathCoding
  • LLM evaluation pipelines often have many candidate judges: general LLM-as-a-judge prompts, reward models, safety classifiers, confidence variants, and task-specific verifiers.
  • The deployment question is not only which judge is best, but which judges should be called, on which examples, and when panel construction should stop.
Open paper
Citations: 0

Match reason: Matches selected tags (Math).

Score: 65% Moderate protocol signal Freshness: Hot Status: Fallback
Automatic Metrics Long Horizon Math
  • Customer-service LLM agents must follow organizational policy when acting on a user's behalf.
  • Runtime safeguards can intervene on risky actions, but action-local checks do not guide an agent through a multi-step procedure.
Open paper

Match reason: Matches selected tags (Math).

Score: 65% High protocol signal Freshness: Hot Status: Fallback
Automatic Metrics Long Horizon Math
  • For instance, LoRA-GA^2 surpasses the leading baseline by an average of 0.66 points on the GLUE benchmark, and outperforms the strongest baseline by 1.03 points on GSM8K and 0.87 points on HumanEval, respectively.
Open paper
Right in the Right Way: LM Training with Verifiable Rewards and Human Demonstrations

Mehul Damani, Isha Puri, Idan Shenfeld, Jacob Andreas · Jul 1, 2026

Citations: 0

Match reason: Matches selected tags (Math).

Score: 58% Moderate protocol signal Freshness: Warm Status: Ready
Demonstrations Automatic Metrics MathCoding
  • We propose an adversarial generator-discriminator framework that augments verifiable rewards with a learned signal from human demonstrations.
  • In story generation, our method significantly improves win rate while producing stories that are diverse and more human-like.
Open paper
SABER-Math: Automated Benchmark for Information Retrieval Evaluation in Mathematics

Nikolay Georgiev, Maria Drencheva, Kseniia Ibragimova, Ivo Petrov, Dimitar I. Dimitrov, Martin Vechev · Jun 29, 2026

Citations: 0

Match reason: Matches selected tags (Math).

Score: 58% Moderate protocol signal Freshness: Warm Status: Ready
Pairwise Preference Automatic Metrics Math
  • As agentic AI systems tackle more complex mathematical tasks, they increasingly rely on information retrieval (IR) to search problem databases, theorem libraries, and educational resources.
  • Importantly, we show that general-purpose IR benchmarks such as MTEB do not reliably predict mathematical performance, especially for recent embedding models, highlighting the need for math-specific retrieval benchmarks.
Open paper
Citations: 0

Match reason: Matches selected tags (Math).

Score: 58% High protocol signal Freshness: Warm Status: Ready
Pairwise Preference Automatic Metrics Math
  • Across seven models and three mathematical reasoning benchmarks (GSM1K, MATH500, AIME 2025), cliff tokens act as failure triggers; deleting the first cliff token and resampling recovers pass@64 to 1.0, while keeping it limits recovery to…
  • Trained on GSM8K, Cliff-DPO improves accuracy across benchmarks by up to +6.6.
Open paper

Match reason: Matches selected tags (Math).

Score: 58% High protocol signal Freshness: Warm Status: Ready
Rubric Rating Automatic Metrics Math
  • We propose a lightweight diagnostic based on the gap between solving-oriented and pedagogy-oriented benchmark performance.
  • Using public MathTutorBench leaderboard results, we show that these dimensions are only partially aligned: across eight publicly reported models, the correlation between solving and pedagogy composites is 0.421, and several models shift…
Open paper
SEAL: Can Saturated Benchmarks Be Revived by LLM-as-a-Meta-Judge?

Jiamin Chen, Yidi Wu, Qiexiang Wang, Qianben Chen, Yuchen Li, Yansen Zhang · May 28, 2026

Citations: 0

Match reason: Matches selected tags (Math).

Score: 58% Moderate protocol signal Freshness: Warm Status: Ready
Pairwise Preference Automatic Metrics MathCoding
  • Therefore, we present Seeded Elimination with Adaptive LLM-as-a-Meta-Judge, a self-improving evaluation protocol for extracting latent ranking signal from saturated benchmarks.
  • We evaluate SEAL on multiple saturated benchmarks covering code generation, mathematical reasoning, knowledge-intensive question answering, and tool-use agent task completion.
Open paper
Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles

Jinyang Wu, Guocheng Zhai, Ruihan Jin, Yuhao Shen, Zhengxi Lu, Fan Zhang · May 21, 2026

Citations: 0

Match reason: Matches selected tags (Math).

Score: 58% Moderate protocol signal Freshness: Warm Status: Ready
Expert Verification Automatic Metrics MathCoding
  • In this paper, we present Maestro (Multimodal Agent for Expert-Skill Targeted Reinforced Orchestration), a Reinforcement Learning (RL)-driven orchestration framework that reframes heterogeneous multimodal tasks as a sequential…
  • We evaluate Maestro across ten representative multimodal benchmarks spanning mathematical reasoning, chart understanding, high-resolution perception, and domain-specific analysis.
Open paper
Efficient and Trainable Language Model Test-Time Scaling via Local Branch Routing

Yutong Yin, Mingyu Jin, Jin Pan, Changyi Yang, Zijie Xia, Dhruv Pai · Jun 24, 2026

Citations: 0

Match reason: Matches selected tags (Math).

Score: 58% Moderate protocol signal Freshness: Warm Status: Fallback
Automatic Metrics Long Horizon Math
  • On mathematical reasoning benchmarks, LBR improves both Pass@1 and Pass@32 over discrete chain-of-thought, vanilla discrete-token RLVR, and RL-compatible soft-token branching baselines.
Open paper
Blockwise Policy-Drift Gating for On-Policy Distillation

Liwen Zheng, Haiyun Jiang · Jun 23, 2026

Citations: 0

Match reason: Matches selected tags (Math).

Score: 58% High protocol signal Freshness: Warm Status: Fallback
Automatic Metrics Long Horizon Math
  • In a six-variant Qwen3 math reasoning benchmark with a uniform 200-step training budget for all trained variants, we use pass@8 as the primary problem-level solve-rate metric.
  • On Teacher-TopK/LSM, Block64 gives the best four-benchmark mean pass@8 among trained students.
Open paper
Scaling Laws for Agent Harnesses via Effective Feedback Compute

Xuanliang Zhang, Dingzirui Wang, Keyan Xu, Qingfu Zhu, Wanxiang Che · May 28, 2026

Citations: 0

Match reason: Matches selected tags (Math).

Score: 58% Moderate protocol signal Freshness: Warm Status: Fallback
Automatic Metrics Tool Use MathLaw
  • Agent harnesses shape language-model performance by controlling tool use, feedback, verification, memory, and repair.
  • Across synthetic, real, held-out, and prospective evaluations, EFC-based coordinates outperform raw-compute baselines and SAS.
Open paper
Index SLM Technical Report

Tianjiao Li, Lusheng Zhang, Shien He, Xiaojing Liu, Tianxing Yan, Mengran Yu · Jul 10, 2026

Citations: 0

Match reason: Matches selected tags (Math).

Score: 52% Sparse protocol signal Freshness: Warm Status: Fallback
Pairwise Preference MathCoding
  • The series comprises four models: Index-1.9B-Base, a foundation model with 1.9 billion non-embedding parameters pre-trained on 2.8 trillion predominantly Chinese and English tokens; Index-1.9B-Pure, a control variant trained with an…
  • On a suite of standard benchmarks covering examination, reasoning, mathematics, and code, Index-1.9B-Base attains an average score of 64.92, competitive with or exceeding open models of several times its size.
Open paper
Online Safety Monitoring for LLMs

Mona Schirmer, Metod Jazbec, Alexander Timans, Christian Naesseth, Maja Waldron, Eric Nalisnick · Jul 2, 2026

Citations: 0

Match reason: Matches selected tags (Math).

Score: 52% Sparse protocol signal Freshness: Warm Status: Fallback
Red Team Math
  • Monitoring outputs online and raising an alarm when safety can no longer be assumed is therefore critical.
Open paper
Are We Measuring Strategy or Phrasing? The Gap Between Surface- and Approach-Level Diversity in LLM Math Reasoning

Sangmook Lee, Minbeom Kim, Jeonghye Kim, Dohyung Kim, Sojeong Rhee, Kyomin Jung · Jun 29, 2026

Citations: 0

Match reason: Matches selected tags (Math).

Score: 52% Sparse protocol signal Freshness: Warm Status: Fallback
Pairwise Preference Math
  • Using a human-calibrated LLM judge framework, we show that prior diversity measures are unreliable proxies for approach-level diversity, and this mismatch carries over to diversity-aware RLVR, where target metrics are preserved while…
  • However, optimizing an LLM judge diversity reward during training causes the policy to exploit judge-specific preferences rather than broaden its approaches, leaving direct optimization of approach-level diversity as an open problem.
Open paper
LatentRevise: Learning from Zero-Hit Reasoning

Yiqiu Guo, Xueting Han, Qi Jia, Guangtao Zhai, Jing Bai · Jun 29, 2026

Citations: 0

Match reason: Matches selected tags (Math).

Score: 52% Sparse protocol signal Freshness: Warm Status: Fallback
Critique Edit Math
  • Used as training data, these trajectories improve SFT and RLVR on math benchmarks over standard baselines.
Open paper