Skip to content
OpenTrain AIFor AI Companies

Researcher Tools

Human Feedback and Eval Paper Explorer

A focused feed for RLHF, preference data, rater protocols, agent evaluation, and LLM-as-judge research. Every paper includes structured metadata for quick triage.

Total papers: 169 Search mode: keyword Shortlist (0) RSS

Featured Papers

Popular high-signal papers with direct links to full protocol pages.

Weekly Eval Paper Digest

The top RLHF, evaluation, and human feedback papers — curated and summarized every Friday.

No spam. Unsubscribe anytime.

Start Here By Objective

Pick your immediate research objective and jump directly to high-signal pages, not generic search.

Scale Your Evaluation Team

Need human evaluators for your benchmark or preference study? OpenTrain sources pre-vetted domain experts into your annotation pipeline.

Match reason: Matches selected tags (Medicine).

Score: 58% Moderate protocol signal Freshness: Warm Status: Ready
Expert Verification Automatic Metrics Medicine
  • We introduce a reusable pipeline for generating terminology-grounded HL7 FHIR R4 bundles from unstructured text, enabling controllable evaluation of clinical decision support systems over structured inputs.
  • Evaluation on MedCase-Structured reveals consistently lower diagnostic accuracy for LLMs on structured FHIR inputs than with plain text, highlighting the importance of deployment-aligned benchmarking.
Open paper
DirectorBench: Diagnosing Long-Form Video Generation with Personalized Multi-Agent Evaluation

Jiamin Chen, Qianben Chen, Jiawen Zhang, Yidi Wu, Yuchen Li, Xiaokun Zhang · May 28, 2026

Citations: 0

Match reason: Matches selected tags (Medicine).

Score: 58% High protocol signal Freshness: Warm Status: Ready
Pairwise Preference Human Eval Multi Agent Medicine
  • However, evaluating such videos remains challenging, since existing benchmarks largely focus on local visual quality, short-horizon temporal consistency, or generic prompt alignment, and provide limited diagnosis of workflow failures and…
  • We introduce DirectorBench, a personalized multi-agent diagnostic benchmark for long-form video generation.
Open paper
Citations: 0

Match reason: Matches selected tags (Medicine).

Score: 58% High protocol signal Freshness: Warm Status: Fallback
Automatic Metrics Tool Use MedicineCoding
  • We introduce ToolBench-X, a benchmark for evaluating agents under recoverable reliability hazards.
  • These results suggest that tool-use evaluation should move beyond function-call accuracy toward task completion under unreliable tool environments.
Open paper
SHERLOC: Structured Diagnostic Localization for Code Repair Agents

Hovhannes Tamoyan, Sean Narenthiran, Erik Arakelyan, Mira Mezini, Boris Ginsburg · Jun 23, 2026

Citations: 0

Match reason: Matches selected tags (Medicine).

Score: 58% High protocol signal Freshness: Warm Status: Fallback
Automatic Metrics Tool Use MedicineCoding
  • We introduce SHERLOC (Structured Hypothesis-driven Exploration and Reasoning for Localization), a training-free framework pairing a reasoning LLM with compact repository tools and self-recovery, without fine-tuning or multi-agent…
  • SHERLOC reaches state-of-the-art localization across model scales: 84.33% accuracy@1 on SWE-Bench Lite and 81.27% recall@1 on SWE-Bench Verified; at ~30B parameters, it matches or outperforms other agentic methods.
Open paper
MedBench v5: A Dynamic, Process-Oriented, and Hallucination-Aware Benchmark for Clinical Multimodal Models

Jinru Ding, Chuchu Jiang, Lu Lu, Wenrao Pang, Mouxiao Bian, Zhuangzhi Gao · Jun 23, 2026

Citations: 0

Match reason: Matches selected tags (Medicine).

Score: 58% Moderate protocol signal Freshness: Warm Status: Fallback
Simulation Env Long Horizon Medicine
  • Existing medical AI benchmarks lack process visibility, atomic skill evaluation, and integrated hallucination detection.
  • We introduce MedBench v5, a redesigned benchmark for clinical multimodal models (language, vision-language, and agent systems) that moves from static QA to dynamic, process-oriented evaluation.
Open paper
EvoRubric: Self-Evolving Rubric-Driven RL for Open-Ended Generation

Xin Guan, Xiaomeng Hu, Shen Huang, Zhenyi Wang, Bo Zhang, Zijian Li · May 28, 2026

Citations: 0

Match reason: Matches selected tags (Medicine).

Score: 55% Moderate protocol signal Freshness: Warm Status: Fallback
Rubric Rating Medicine
  • Current rubric-based RL methods mitigate this by employing explicit criteria; however, they rely heavily on static, human-annotated rubrics that inevitably cause policy lag, or expensive external proprietary models for dynamic updates.
  • Notably, our framework is compatible with human-expert priors.
Open paper
Self-Prompting Small Language Models for Privacy-Sensitive Clinical Information Extraction

Yao-Shun Chuang, Tushti Mody, Uday Pratap Singh, Shirindokht Shiraz, Chun-Teh Lee, Ryan Brandon · May 5, 2026

Citations: 0

Match reason: Matches selected tags (Medicine).

Score: 53% Moderate protocol signal Freshness: Cold Status: Ready
Pairwise Preference Automatic Metrics Medicine
  • Using 1,200 annotated notes, we evaluated candidate open-weight models with multi-prompt ensemble inference and further adapted selected models using QLoRA-based supervised fine-tuning and direct preference optimization.
  • Model performance varied substantially, highlighting the need for task-specific evaluation rather than reliance on generic benchmarks.
Open paper
Safety and accuracy follow different scaling laws in clinical large language models

Sebastian Wind, Tri-Thien Nguyen, Jeta Sopa, Mahshad Lotfinia, Sebastian Bickelhaup, Michael Uder · May 5, 2026

Citations: 0

Match reason: Matches selected tags (Medicine).

Score: 53% Moderate protocol signal Freshness: Cold Status: Ready
Expert Verification Automatic Metrics LawMedicine
  • We introduce SaFE-Scale, a framework for measuring how clinical LLM safety changes across model scale, evidence quality, retrieval strategy, context exposure, and inference-time compute.
  • To instantiate this framework, we introduce RadSaFE-200, a Radiology Safety-Focused Evaluation benchmark of 200 multiple-choice questions with clinician-defined clean evidence, conflict evidence, and option-level labels for high-risk error,…
Open paper
MedMosaic: A Challenging Large Scale Benchmark of Diverse Medical Audio

Harshit Rajgarhia, Shuubham Ojha, Asif Shaik, Akhil Pothanapalli, Rachuri Lokesh, Abhishek Mukherji · May 1, 2026

Citations: 0

Match reason: Matches selected tags (Medicine).

Score: 53% Moderate protocol signal Freshness: Cold Status: Ready
Expert Verification Automatic Metrics Medicine
  • Thus, existing benchmarks tend to underrepresent complex medical audio scenarios.
  • To address this challenge, we present MedMosaic, a medical audio question-answering dataset designed to benchmark language and audio reasoning models under realistic clinical constraints.
Open paper
PolicyAlign: Direct Policy-Based Safety Alignment for Large Language Models

Chang Wu, Junfeng Fang, Houcheng Jiang, Kai Tang, Pengyu Cheng, Xiaoxi Jiang · Jun 24, 2026

Citations: 0

Match reason: Matches selected tags (Medicine).

Score: 52% Sparse protocol signal Freshness: Warm Status: Fallback
Pairwise PreferenceDemonstrations LawMedicine
  • Safety alignment of large language models (LLMs) typically depends on high-quality supervision data, such as safe demonstrations or preference pairs.
  • To address this, we propose PolicyAlign, a simple yet effective framework for directly aligning LLMs with safety policies.
Open paper

Match reason: Matches selected tags (Medicine).

Score: 52% Sparse protocol signal Freshness: Warm Status: Fallback
Expert Verification Medicine
  • The framework coordinates specialized agents for clinical text, longitudinal EHR, medical imaging, physiological sensor signals, guideline retrieval, uncertainty auditing, and referral planning.
  • We also outline a real-data evaluation design using public and credentialed clinical datasets spanning EHR, radiology, ECG, ICU time series, and referral-proxy outcomes.
Open paper
Before the Labels: How Dataset Construction Shapes Suicidality Detection in Clinical Text

Priyanshi Garg, Ishita Rao, Jieqiong Ding, Amandalynne Paullada · Jun 17, 2026

Citations: 0

Match reason: Matches selected tags (Medicine).

Score: 52% Sparse protocol signal Freshness: Warm Status: Fallback
Expert Verification Medicine
  • We show how governance constraints, ICD-based cohort selection, single-annotator labeling, and hospital-stay-level aggregation produce labels that reflect clinician-documented judgments, treat suicidality as a bounded episode, and assume…
Open paper
eCream-MedCorpus A Large-Scale Corpus of Clinical Notes for Italian

Tiziano Labruna, Guido Bertolini, Pietro Ferrazzi, Bernardo Magnini · Jun 10, 2026

Citations: 0

Match reason: Matches selected tags (Medicine).

Score: 52% Sparse protocol signal Freshness: Warm Status: Fallback
Expert VerificationCritique Edit Medicine
  • Finally, we propose CRF-filling as a novel structured information extraction benchmark, and provide zero-shot baseline resulting from Gemma-27B and MedGemma-27B.
Open paper
Leveraging Social Media Data for COVID-19 Studies

Nur Hafieza Ismail, Nur Shazwani Kamarudin, Nurol Husna Che Rose · Jun 9, 2026

Citations: 0

Match reason: Matches selected tags (Medicine).

Score: 52% Sparse protocol signal Freshness: Warm Status: Fallback
Expert Verification Medicine
Open paper
CRITIC-R1: Learning Structured Critics for Retrieval-Augmented Generation

Wenhan Xiao, Ziwei Zhang, Chuanyue Yu, Xingcheng Fu, Qingyun Sun, Runhua Xu · May 28, 2026

Citations: 0

Match reason: Matches selected tags (Medicine).

Score: 52% Sparse protocol signal Freshness: Warm Status: Fallback
Critique Edit MedicineCoding
  • To learn these capabilities, we design two reward functions: Conservative Judgement Alignment (CJA) first encourages calibrated high-level judgements while mitigating the over-aggressive phenomenon, whereas Diagnostic Quality Alignment…
  • Experiments across five QA benchmarks show that CRITIC-R1 consistently improves answer quality over strong RAG baselines.
Open paper
Holistic Evaluation and Failure Diagnosis of AI Agents

Netta Madvil, Gilad Dym, Alon Mecilati, Edo Dekel, Jonatan Liberman, Rotem Brazilay · May 14, 2026

Citations: 0

Match reason: Matches selected tags (Medicine).

Score: 53% High protocol signal Freshness: Cold Status: Fallback
Automatic Metrics Long Horizon Medicine
  • We present a holistic agent evaluation framework that pairs top-down agent-level diagnosis with bottom-up span-level evaluation, decomposing analysis into independent per-span assessments.
  • On the TRAIL benchmark, our framework achieves state-of-the-art results across all metrics on both GAIA and SWE-Bench, with relative gains over the strongest prior baselines of up to 38% on category F1, up to 3.5x on localization accuracy,…
Open paper
Citations: 0

Match reason: Matches selected tags (Medicine).

Score: 53% High protocol signal Freshness: Cold Status: Fallback
Automatic Metrics Multi Agent Medicine
  • To address this, we present CuraView, a multi-agent framework for sentence-level detection and evidence-grounded explanation of faithfulness hallucinations in discharge summaries.
  • We evaluate CuraView on a subset of 250 patients from the Discharge-Me benchmark, with 50 patients held out for testing.
Open paper
Detecting Stealth Sycophancy in Mental-Health Dialogue with Dynamic Emotional Signature Graphs

Tianze Han, Beining Xu, Hanbo Zhang, Yongming Lu · May 5, 2026

Citations: 0

Match reason: Matches selected tags (Medicine).

Score: 53% High protocol signal Freshness: Cold Status: Fallback
Automatic Metrics Long Horizon Medicine
  • As conversational AI therapists are increasingly used in psychological support settings, reliable offline evaluation of therapeutic response quality remains an open problem.
  • We evaluate DESG on a constructed diagnostic stress-test benchmark of 3{,}000 dialogue windows from EmpatheticDialogues, ESConv, and CRADLE-Dialogue, covering peer support, counseling dialogue, and crisis-oriented interaction.
Open paper