Skip to content
OpenTrain AIFor AI Companies

Researcher Tools

Human Feedback and Eval Paper Explorer

A focused feed for RLHF, preference data, rater protocols, agent evaluation, and LLM-as-judge research. Every paper includes structured metadata for quick triage.

Total papers: 501 Search mode: keyword Shortlist (0) RSS

Featured Papers

Popular high-signal papers with direct links to full protocol pages.

Weekly Eval Paper Digest

The top RLHF, evaluation, and human feedback papers — curated and summarized every Friday.

No spam. Unsubscribe anytime.

Start Here By Objective

Pick your immediate research objective and jump directly to high-signal pages, not generic search.

Scale Your Evaluation Team

Need human evaluators for your benchmark or preference study? OpenTrain sources pre-vetted domain experts into your annotation pipeline.

No exact ID match for "2608.28833" yet. Showing current high-signal papers so you can continue browsing while this paper is indexed.
PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization

Boryeong Cho, Sumyeong Ahn, Se-Young Yun · Aug 31, 2026

Citations: 0

Match reason: Matched by broad semantic/index fallback.

Score: 45% Moderate protocol signal Freshness: Hot Status: Ready
Pairwise Preference Automatic Metrics General
  • To address this, we propose Posterior Label Correction DPO (PLC-DPO) to robustly optimize preferences by routing each pair's training signal as a clean, flip, or tie case.
  • Across 57 dataset-model-benchmark cells, PLC-DPO obtains the best mean win rate against DPO (60.5 vs.
Open paper
Collapsibility of Performance Metrics in Clinical Predictive AI

João Matos, Ben Van Calster, Richard D. Riley, Paula Dhiman, Gary S. Collins · Aug 31, 2026

Citations: 0

Match reason: Matched by broad semantic/index fallback.

Score: 45% Moderate protocol signal Freshness: Hot Status: Ready
Automatic Metrics MathMedicine
  • Fairness evaluations commonly rely on performance analyses across subgroups.
  • Conclusions: Non-collapsibility of performance metrics has important consequences for reporting, model appraisal, and fairness evaluation.
Open paper
BiG-SURE - Bipartite Graph for Semantic Uncertainty and Reliability Estimation of LLMs

Debarpan Bhattacharya, Malay Phadke, Sriram Ganapathy · Aug 31, 2026

Citations: 0

Match reason: Matched by broad semantic/index fallback.

Score: 42% Moderate protocol signal Freshness: Hot Status: Ready
Automatic Metrics Multilingual
  • Reliable uncertainty estimation is a crucial requirement for deploying large language models (LLMs) and vision-language models (VLMs) in safety-critical settings, especially when the model parameters are not accessible (black-box).
Open paper
Cost-efficient Active Learning for Referring Image Segmentation and Grounding

Junbeom Hong, Seonghoon Yu, Hyung Rok Jung, Sundong Kim, Jeany Son · Aug 31, 2026

Citations: 0

Match reason: Matched by broad semantic/index fallback.

Score: 42% Moderate protocol signal Freshness: Hot Status: Ready
Automatic Metrics General
  • Collecting natural-language referring expressions along with region annotations, such as masks or boxes, is a major bottleneck in visual grounding (VG), as annotators must write descriptions that distinguish target regions from visually…
  • We also design a referring-expression annotation interface that helps annotators quickly focus on writing discriminative language with a few clicks.
Open paper
SwarmBench: Can Large Language Models Act as Agent Swarm Orchestrators?

Jinshan Gao, Zhuoran Jin, Tianyi Men, Kang Liu, Jun Zhao · Aug 31, 2026

Citations: 0

Match reason: Matched by broad semantic/index fallback.

Score: 45% High protocol signal Freshness: Hot Status: Fallback
Automatic Metrics Multi Agent General
  • Large language model-based multi-agent systems are evolving from fixed interaction topologies toward dynamically orchestrated Agent Swarms.
  • We propose SwarmBench, a benchmark that evaluates model performance from multiple perspectives, including accuracy, efficiency, cost, and process quality.
Open paper
TaxCE : A Framework for Automated Taxonomy Construction and Evaluation at Scale

Sandeep Sricharan Mukku, Albert Aristotle Nanda, Rohit Pyati · Aug 31, 2026

Citations: 0

Match reason: Matched by broad semantic/index fallback.

Score: 38% Sparse protocol signal Freshness: Hot Status: Ready
Human Eval General
  • Existing approaches either produce shallow hierarchies, neglect long-tail topics, or lack rigorous evaluation frameworks.
  • We also introduce three corpus-grounded evaluation metrics, Exclusivity, Exhaustivity, and Granularity (EEG), and integrate them into a metrics-in-the-loop iterative refinement mechanism that diagnoses deficiencies and applies targeted…
Open paper
MURANO: Design, Run, and Reproduce Mechanistic Interpretability Experiments as Composable Pipelines

Alireza Bayat Makou, Emirhan Böge, Phu Gia Hoang, Federico Tiblias, Jingcheng Niu, Subhabrata Dutta · Aug 31, 2026

Citations: 0

Match reason: Matched by broad semantic/index fallback.

Score: 35% Sparse protocol signal Freshness: Hot Status: Ready
General
  • These studies often combine loading, recording, attribution, intervention, and evaluation, while existing libraries tend to focus on different parts of this workflow.
Open paper
Fine-Grained Multi Image Object Hallucination Benchmark

Joonki Min, Chaeyun Kim, Hyungwook Choi, Yejin Kim, Kihyun Kim, Yohan Jo · Aug 31, 2026

Citations: 0

Match reason: Matched by broad semantic/index fallback.

Score: 35% Sparse protocol signal Freshness: Hot Status: Ready
General
  • Existing benchmarks, designed primarily for single-image settings or providing only high-level multi-image assessments, cannot systematically diagnose how visual complexity and reasoning demands trigger hallucination.
  • To address this gap, we introduce MIOH, a fine-grained multi-image object hallucination benchmark that systematically evaluates object hallucination across four foundational tasks (existence, counting, attribute, position) through three…
Open paper
Citations: 0

Match reason: Matched by broad semantic/index fallback.

Score: 35% Sparse protocol signal Freshness: Hot Status: Ready
General
  • To evaluate the adaptation efficacy, we also construct a novel domain-specific benchmark that covers six editorial tasks.
  • Crucially, we demonstrate the importance of targeted evaluation in the adaptation process, as an existing Swedish benchmark largely fails to capture the models' in-domain performance gains.
Open paper
DiffSAC: Diffusion-guided Sampling for Consensus-based Robust Estimation

Chang Nie, Guangming Wang, Zhe Liu, Hesheng Wang · Aug 31, 2026

Citations: 0

Match reason: Matched by broad semantic/index fallback.

Score: 35% Sparse protocol signal Freshness: Hot Status: Ready
General
  • However, traditional methods suffer from inefficient sampling as they struggle to identify effective minimum sets before hypothesis evaluation.
  • Consequently, DiffSAC outputs a small number of high-quality minimum sets, enabling identification of the best hypothesis via consensus evaluation.
Open paper