Skip to content
OpenTrain AIFor AI Companies

Researcher Tools

Human Feedback and Eval Paper Explorer

A focused feed for RLHF, preference data, rater protocols, agent evaluation, and LLM-as-judge research. Every paper includes structured metadata for quick triage.

Total papers: 95 Search mode: keyword Shortlist (0) RSS

Featured Papers

Popular high-signal papers with direct links to full protocol pages.

Weekly Eval Paper Digest

The top RLHF, evaluation, and human feedback papers — curated and summarized every Friday.

No spam. Unsubscribe anytime.

Start Here By Objective

Pick your immediate research objective and jump directly to high-signal pages, not generic search.

Scale Your Evaluation Team

Need human evaluators for your benchmark or preference study? OpenTrain sources pre-vetted domain experts into your annotation pipeline.

Match reason: Matches selected tags (Demonstrations).

Score: 65% High protocol signal Freshness: Hot Status: Ready
Demonstrations Automatic Metrics General
  • On the full GPQA Diamond benchmark (198 graduate-level science questions), majority voting reduces per-problem accuracy on a majority of problems for two instruction-tuned models from different families: 56.6% of problems for Qwen2.5-7B and…
Open paper
OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories

Changhao Xiang, Shilin Zhang, Zheng Ma, Kanzhi Cheng, Ruize Ma, Yi Feng · Aug 9, 2026

Citations: 0

Match reason: Matches selected tags (Demonstrations).

Score: 65% Moderate protocol signal Freshness: Hot Status: Ready
Demonstrations Tool Use Law
  • Visual tool use has emerged as a fundamental capability for multimodal agents to actively acquire evidence beyond a fixed image encoding.
  • Using this framework, we construct OpenVisTool-42K, a dataset spanning five visual reasoning domains, together with OpenVisTool-Bench, a benchmark covering the same domains.
Open paper
Right in the Right Way: LM Training with Verifiable Rewards and Human Demonstrations

Mehul Damani, Isha Puri, Idan Shenfeld, Jacob Andreas · Jul 1, 2026

Citations: 0

Match reason: Matches selected tags (Demonstrations).

Score: 58% Moderate protocol signal Freshness: Warm Status: Ready
Demonstrations Automatic Metrics MathCoding
  • We propose an adversarial generator-discriminator framework that augments verifiable rewards with a learned signal from human demonstrations.
  • In story generation, our method significantly improves win rate while producing stories that are diverse and more human-like.
Open paper

Match reason: Matches selected tags (Demonstrations).

Score: 58% High protocol signal Freshness: Warm Status: Ready
Demonstrations Automatic Metrics General
  • Agentic search equips large language models with dynamic retrieval abilities, but existing reinforcement learning methods remain limited by reward sparsity in knowledge boundary calibration -- deciding when to trust parametric memory, when…
  • Experiments on multiple benchmarks show that KbSD consistently improves both task accuracy and hallucination mitigation over strong baselines, with the largest gains appearing in the challenging quadrants where sparse rewards are least…
Open paper
Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments

Qiuyue Wang, Mingsheng Li, Jian Guan, Jinhui Ye, Sicheng Xie, Yitao Liu · May 28, 2026

Citations: 0

Match reason: Matches selected tags (Demonstrations).

Score: 58% Moderate protocol signal Freshness: Warm Status: Ready
Demonstrations Simulation Env Long Horizon General
  • Qwen-VLA is trained with a large-scale joint pretraining recipe over diverse data sources, including robotics manipulation trajectories, human egocentric demonstrations, synthetic simulation data, vision-and-language navigation data,…
  • Experiments on manipulation, navigation, and trajectory-centric benchmarks show consistent multi-task performance and out-of-distribution generalization under variations in scene layout, background, lighting, object configuration, and robot…
Open paper

Match reason: Matches selected tags (Demonstrations).

Score: 58% Moderate protocol signal Freshness: Warm Status: Ready
Demonstrations Automatic Metrics Law
  • We benchmark seven models from five providers on 273 validated court decisions from Ukraine's state registry (EDRSR), measuring tokenizer fertility and zero-shot performance on three tasks.
  • To support reproducibility and address the absence of Ukrainian from legal NLP benchmarks, we release a public dataset of 14,452 court decisions spanning 2008-2026, annotated with seven outcome labels across three temporal epochs that…
Open paper
GRaSp: Automatic Example Optimization for In-Context Learning in Low-Data Tasks

Simen Bihaug-Frøyland, Henrik Brådland · May 8, 2026

Citations: 0

Match reason: Matches selected tags (Demonstrations).

Score: 58% Moderate protocol signal Freshness: Warm Status: Ready
Demonstrations Automatic Metrics General
  • We evaluate GRaSp on financial named entity recognition (FiNER-139), comparing synthetic and human-annotated candidate pools across pool sizes of 500 and 5000.
Open paper
ContraFix: Skill-Enhanced Contrastive Runtime Analysis for Vulnerability Repair

Simiao Liu, Fang Liu, Peiding Wang, Taichuan Li, Yinghao Zhu, Xiaoli Lian · May 17, 2026

Citations: 0

Match reason: Matches selected tags (Demonstrations).

Score: 55% Moderate protocol signal Freshness: Warm Status: Fallback
Demonstrations General
  • We present ContraFix, an agentic AVR framework that constructs such evidence through contrastive runtime analysis.
  • A semantic audit of benchmark-validated SEC-Bench patches shows that 58.2% of ContraFix's patches are semantically correct, compared with 31.3% for the strongest baseline, indicating that the proposed framework improves semantic correctness…
Open paper
Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL

Sudong Wang, Weiquan Huang, Xiaomin Yu, Zuhao Yang, Hehai Lin, Keming Wu · Apr 30, 2026

Citations: 0

Match reason: Matches selected tags (Demonstrations).

Score: 53% Moderate protocol signal Freshness: Cold Status: Ready
Demonstrations Automatic Metrics Coding
  • Experiments on Qwen3-VL show that PRISM consistently improves downstream RLVR performance across multiple RL algorithms (GRPO, DAPO, GSPO) and diverse multimodal benchmarks, improving average accuracy by +4.4 and +6.0 points over the…
Open paper

Match reason: Matches selected tags (Demonstrations).

Score: 52% Sparse protocol signal Freshness: Warm Status: Fallback
Demonstrations General
  • LLM agents carry conclusions across steps and sessions in compressed memory, and memory products (e.g., mem0, LangMem) rewrite conversation into stored "facts" that later steps trust.
  • We show this rewriting manufactures confidence: across our constructed agent settings, a casual, hedged remark becomes a confident, dated assertion the agent then obeys like a verified fact, granting every above-clearance request it faces.
Open paper
PolicyAlign: Direct Policy-Based Safety Alignment for Large Language Models

Chang Wu, Junfeng Fang, Houcheng Jiang, Kai Tang, Pengyu Cheng, Xiaoxi Jiang · Jun 24, 2026

Citations: 0

Match reason: Matches selected tags (Demonstrations).

Score: 52% Sparse protocol signal Freshness: Warm Status: Fallback
Pairwise PreferenceDemonstrations LawMedicine
  • Safety alignment of large language models (LLMs) typically depends on high-quality supervision data, such as safe demonstrations or preference pairs.
  • To address this, we propose PolicyAlign, a simple yet effective framework for directly aligning LLMs with safety policies.
Open paper
Learning-augmented robotic automation for real-world manufacturing

Yunho Kim, Quan Nguyen, Taewhan Kim, Youngjin Heo, Joonho Lee · Apr 24, 2026

Citations: 0

Match reason: Matches selected tags (Demonstrations).

Score: 46% Sparse protocol signal Freshness: Cold Status: Fallback
Demonstrations General
  • Here we present Learning-Augmented Robotic Automation, a hybrid system that integrates learned task controllers and a neural 3D safety monitor into conventional industrial workflows.
  • We deployed the system on an electric-motor production line to automate deformable cable insertion and soldering under real manufacturing constraints, a step previously performed manually by human workers.
Open paper
Removing Sandbagging in LLMs by Training with Weak Supervision

Emil Ryd, Henning Bartsch, Julian Stastny, Joe Benton, Vivek Hebbar · Apr 23, 2026

Citations: 0

Match reason: Matches selected tags (Demonstrations).

Score: 46% Sparse protocol signal Freshness: Cold Status: Fallback
Demonstrations MathCoding
  • As AI systems begin to automate complex tasks, supervision increasingly relies on weaker models or limited human oversight that cannot fully verify output quality.
Open paper

Daily Archives