Skip to content
OpenTrain AIFor AI Companies

Researcher Tools

Human Feedback and Eval Paper Explorer

A focused feed for RLHF, preference data, rater protocols, agent evaluation, and LLM-as-judge research. Every paper includes structured metadata for quick triage.

Total papers: 501 Search mode: keyword Shortlist (0) RSS

Featured Papers

Popular high-signal papers with direct links to full protocol pages.

Weekly Eval Paper Digest

The top RLHF, evaluation, and human feedback papers — curated and summarized every Friday.

No spam. Unsubscribe anytime.

Start Here By Objective

Pick your immediate research objective and jump directly to high-signal pages, not generic search.

Scale Your Evaluation Team

Need human evaluators for your benchmark or preference study? OpenTrain sources pre-vetted domain experts into your annotation pipeline.

Citations: 0

Match reason: Matches selected tags (General).

Score: 65% High protocol signal Freshness: Hot Status: Ready
Pairwise Preference Llm As JudgeAutomatic Metrics General
  • Personalized text generation aims to make LLMs write in a specific individual's style, yet existing benchmarks measure task accuracy or preference alignment rather than whether the model's output actually resembles the target author's…
  • We introduce PersonalBench, a benchmark that evaluates inference-time personalization methods through three independent lenses: LUAR (a trained authorship verification model), an LLM-as-judge, and automated stylometrics.
Open paper
The Asymmetric Harms of LLM Compression

Yuan Wu, Mairui Li, Lesia Semenova, Chudi Zhong · Aug 20, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 65% Moderate protocol signal Freshness: Hot Status: Ready
Pairwise Preference Automatic Metrics General
  • Finally, we demonstrate that stable aggregate bias scores can conceal substantial, opposing shifts in stereotypical preferences across demographic subgroups.
  • Together, these findings reveal asymmetric behavioral changes that aggregate performance measures fail to capture, highlighting the need for granular evaluation of compressed models before deployment.
Open paper
Mint-Agent: Introducing Finance-Native Agentic Foundation Models

Mint-Agent Team, Kun Wang, Gavin Zhang, Yaze Geng, Lei Tang, Yaoyang Yi · Aug 17, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 65% High protocol signal Freshness: Hot Status: Ready
Expert Verification Automatic Metrics Long Horizon General
  • We present Mint-Agent, a family of finance-native agentic models designed around these two scales of financial intelligence.
  • Across professional financial benchmarks, our models demonstrate two defining strengths: (1) Reliability: Mint-Ag achieves 98.33% on RFC-Bench, surpassing GPT-5.6-Sol and Claude-Opus-4.8 by 3.66 and 3.00 points; and (2) Executability:…
Open paper
Ask to Be Sure: Informative Interactions for Confident Multi-Turn LLM Recommendation

Cedar Site Bai, Zhenyu Liao, Duanshun Li, Sheikh Sarwar, Huiyuan Chen, Yuan Chen · Aug 16, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 65% Moderate protocol signal Freshness: Hot Status: Ready
Pairwise Preference Automatic Metrics General
  • However, guiding multi-turn interactions to elicit user preferences effectively remains challenging.
  • Existing approaches either use separate reinforcement learning agents with templated interactions or optimize for interactivity judged by another LLM, without measuring how much useful information is actually gained.
Open paper

Match reason: Matches selected tags (General).

Score: 65% High protocol signal Freshness: Hot Status: Fallback
Simulation Env Long Horizon General
  • Audited against causal ground truth from executed replay in a single-agent tool environment (ALFWorld), none of the step-level credit signals used to train LLM agents -- LLM-judge scores, outcome-conditioned logprob ratios, or the policy's…
  • A confidence-only router recovers pivotal steps at chance level, but cuts judge cost by 13.1% per turn (14.0% per trajectory).
Open paper
RippleMem: From Isolated Retrieval to Associative Recollection for Long-Term Agent Memory

Jingbo Ji, Lingyi Li, Xilong Cheng, Yuhao Zhou, Wenji Zhang, Yuting Tan · Aug 13, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 58% High protocol signal Freshness: Warm Status: Ready
Llm As JudgeAutomatic Metrics Long Horizon General
  • LLM-based agents increasingly rely on external memory to support long-horizon reasoning and interaction.
  • Experiments on LoCoMo and LongMemEval-S show that RippleMem achieves the best overall performance across evaluated settings, improving LLM-as-a-Judge accuracy by 3.95% on LoCoMo and up to 11.87% on LongMemEval-S, while reducing graph…
Open paper
LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation

Chenrun Wang, Mingxuan Zhu, Tiancheng Huang, Wenjie Li, Yujie Zhang, Zichen Zhu · Aug 13, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 58% High protocol signal Freshness: Warm Status: Ready
Pairwise Preference Automatic Metrics General
  • To address this challenge, we propose LigBench, an automated evaluation benchmark that enables fine-grained and reliable evaluation of AI research ideas, consistently applicable across different generation distributions.
  • In addition, we introduce PAIR-IQ, a dataset tailored for training pairwise idea judgment models and serving as an auxiliary reference to support more objective comparative evaluation.
Open paper

Match reason: Matches selected tags (General).

Score: 58% High protocol signal Freshness: Warm Status: Ready
Red Team Automatic Metrics Multi Agent General
  • We present a two-tier agentic system that separates a maintained, point-in-time knowledge library from report writing.
  • A portable multi-agent "writer" runtime then composes a contradiction-free, evidence-grounded report at any knowledge cutoff T, reading only evidence with as_of <= T (no look-ahead); red-team verdicts flow back into the librarian.
Open paper

Match reason: Matches selected tags (General).

Score: 58% Moderate protocol signal Freshness: Warm Status: Ready
Pairwise Preference Automatic Metrics General
  • Benchmark contamination is diagnosed today with n-gram overlap, with likelihood-based membership inference, or with canary strings, and each needs something usually unavailable: the training corpus, a well-chosen test statistic, or…
Open paper
PEA-DPO: Perception-Enhanced Alignment Direct Preference Optimization for MLLMs Alignment

Jiawei Feng, Jiancan Wu, Xingyu Zhu, Junkang Wu, Xiang Wang, Xiangnan He · Aug 20, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 58% Sparse protocol signal Freshness: Hot Status: Fallback
Pairwise Preference General
  • Direct Preference Optimization (DPO) has emerged as an effective approach for aligning large language models (LLMs) with human preferences.
  • To address these challenges, we propose Perception-Enhanced Alignment DPO (PEA-DPO), a framework for multimodal LLMs alignment, which explicitly leverages visual preference signals to overcome visual insensitivity.
Open paper
HARP: Hierarchical Adaptive Ranking with Preference-Adaptive Fusion for Query-Based CVE Prioritization

Haochen Liu, Zhengzhang Chen, Haoyu Wang, Yanchi Liu, Jundong Li, Haifeng Chen · Aug 19, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 58% Sparse protocol signal Freshness: Hot Status: Fallback
Pairwise Preference General
  • Vulnerability prioritization is inherently preference dependent, since the same CVE can receive different remediation priority under different operational preference scenarios.
  • In practice, organizations already operate under a preference scenario, but this preference is often implicit and difficult to express as a written prompt instruction, while triage queries usually do not encode it.
Open paper

Match reason: Matches selected tags (General).

Score: 58% High protocol signal Freshness: Warm Status: Fallback
Automatic Metrics Long Horizon General
  • We adapt S2G-RAG's structured sufficiency-and-gap judgment to a frozen Search-R1 pipeline and train a Qwen3.5-2B judge on 3,009 states from 900 disjoint HotpotQA questions.
  • Thus, the trained S2G-style structured judge reduces retrieval while broadly preserving answer accuracy.
Open paper
LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level Consolidation

Dongfang Li, Zixuan Liu, Junmai Wang, Jiahe Huang, Fuhao Li, Bonian Jia · Aug 13, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 58% High protocol signal Freshness: Warm Status: Fallback
Automatic Metrics Long Horizon General
  • Long-horizon LLM agents must preserve information from past interactions to support future tasks.
  • More broadly, our results suggest that the accuracy--cost trade-off of long-term agent memory depends not only on what information is retained, but also on the granularity at which it is consolidated.
Open paper
Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories

Yifei Li, Heng Wang, Lingling Zhang, Muye Huang, Xinyu Zhang, Jiashuai Liu · Aug 13, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 58% High protocol signal Freshness: Warm Status: Fallback
Simulation Env Long Horizon General
  • Retrieval can identify a past trajectory that may matter, yet it does not specify how an acting agent should use that trajectory after users, entities, constraints, or environment state have changed.
  • We identify this post-retrieval reuse step as a distinct bottleneck for long-horizon trajectory memory and formulate an evaluation framework that holds candidate retrieval, target state, model, decoding, and tool budget fixed while varying…
Open paper
ERSkill: Evolving for Skill-Guided Adaptive Memory Retrieval

Haolong Chen, Liang Zhang, Zhuo Li, Lei Xue, Guanrxu Zhu · Aug 13, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 58% Moderate protocol signal Freshness: Warm Status: Fallback
Automatic Metrics Multi Agent General
  • While Large Language Model (LLM) agents increasingly rely on long-term memory for persistent interactions, the retrieval mechanisms governing this memory are rarely treated as evolvable components.
  • Notably, it improves the overall average across F1, BLEU-1, and LLM-judge scores by 31.3\% with Qwen3-Next-80B-A3B-Instruct and by 28.1\% with GPT-5.4-nano.
Open paper

Match reason: Matches selected tags (General).

Score: 58% High protocol signal Freshness: Warm Status: Fallback
Automatic Metrics Long Horizon General
  • We demonstrate two bottlenecks in existing systems: indices built from context-poor captions are unreliable for agentic search, while retrieval ignores a question's temporal intent.
  • To address both bottlenecks, we introduce EgoCITE (Egocentric Context-augmented Indexing and Time-aware Evidence retrieval), a long-horizon agentic memory framework for egocentric QA.
Open paper

Match reason: Matches selected tags (General).

Score: 58% Moderate protocol signal Freshness: Warm Status: Fallback
Automatic Metrics Long Horizon General
  • To address this, we introduce EpicStar, a framework that enables agents to learn memory as policy to tackle long-horizon reasoning.
  • Specifically, the agent maintains a bank of successful past episodes as a heuristic alongside a working memory to track short-term environmental changes.
Open paper
Synthetic Persona Pretraining: Alignment from Token Zero

Julian Minder, Viktor Moskvoretskii, Raghav Singhal, Difan Jiao, Andy Arditi, Shaobo Cui · Aug 13, 2026

Citations: 0

Match reason: Matches selected tags (General).

Score: 52% Sparse protocol signal Freshness: Warm Status: Fallback
Red Team General
  • As language-model-based AI is increasingly deployed in autonomous settings, aligning its goals and values with those of humans becomes critical.
Open paper