HFEPX Hub

CS.CL Papers (Last 30 Days)

Updated from current HFEPX corpus (Apr 27, 2026). 1364 papers are grouped in this hub page.

Read Full Context

Updated from current HFEPX corpus (Apr 27, 2026). 1364 papers are grouped in this hub page. Common evaluation modes: Automatic Metrics, Simulation Env. Most common rater population: Domain Experts. Common annotation unit: Trajectory. Frequent quality control: Calibration. Frequently cited benchmark: DROP. Common metric signal: accuracy. Use this page to compare protocol setup, judge behavior, and labeling design decisions before running new eval experiments. Newest paper in this set is from Mar 31, 2026.

Papers: 1,364 Last published: Mar 31, 2026 Global RSS

Cs.CLLast 30d

Researcher Quick Triage

This hub is best used for protocol triage and replication planning from abstract-level evidence. Quality band: High .

Analysis blocks below are computed from the currently loaded sample (60 of 1,364 total papers in this hub).

All Sampled Papers (60) Replication-Ready Only (17)

High-Signal Coverage

100.0%

60 / 60 sampled papers are not low-signal flagged.

Replication-Ready Set

Benchmark + metric + eval mode explicitly present.

Judge/Human Comparability

Papers containing both `human_eval` and `llm_as_judge`.

17 papers are replication-ready (benchmark + metric + explicit evaluation mode).
0 papers support judge-vs-human agreement analysis.
7 papers report explicit quality controls (calibration/adjudication/IAA).

Primary action: Start with the top 2 papers in “Start Here”, then validate assumptions in the protocol matrix.

Currently showing only replication-ready papers in ranking and matrix sections (17 papers).

Need evaluators for this research workflow?

Post a Job →

Why This Matters For Eval Research

6.6% of papers report explicit human-feedback signals, led by pairwise preferences.
automatic metrics appears in 19.9% of papers in this hub.
DROP is a recurring benchmark anchor for cross-paper comparisons in this page.

Protocol Takeaways

Most common quality-control signal is rater calibration (1.8% of papers).
Rater context is mostly domain experts, and annotation is commonly trajectory-level annotation; use this to scope replication staffing.
Compare papers that report both human_eval and llm_as_judge to quantify judge-human agreement drift.

Benchmark Interpretation

DROP appears in 1.1% of hub papers (15/1364); use this cohort for benchmark-matched comparisons.
MMLU appears in 0.9% of hub papers (12/1364); use this cohort for benchmark-matched comparisons.

Metric Interpretation

accuracy is reported in 13.3% of hub papers (181/1364); compare with a secondary metric before ranking methods.
cost is reported in 5.9% of hub papers (81/1364); compare with a secondary metric before ranking methods.

Researcher Checklist (Expanded)

Researcher Checklist

Gap: Papers with explicit human feedback

Coverage is a replication risk (6.6% vs 45% target).
Gap: Papers reporting quality controls

Coverage is a replication risk (3% vs 30% target).
Gap: Papers naming benchmarks/datasets

Coverage is a replication risk (12.2% vs 35% target).
Moderate: Papers naming evaluation metrics

Coverage is usable but incomplete (30.4% vs 35% target).
Gap: Papers with known rater population

Coverage is a replication risk (6.3% vs 35% target).
Gap: Papers with known annotation unit

Coverage is a replication risk (6.2% vs 35% target).

Strengths

Contains both human-eval and LLM-as-judge protocols for head-to-head methodology comparison.

Known Gaps

Only 3% of papers report quality controls; prioritize calibration/adjudication evidence.
Rater population is under-specified (6.3% coverage).
Annotation unit is under-specified (6.2% coverage).

Suggested Next Analyses

Compare papers that report both human_eval and llm_as_judge to quantify judge-human agreement drift.
Stratify by benchmark (DROP vs MMLU) before comparing methods.
Track metric sensitivity by reporting both accuracy and cost.

Recommended Queries (Expanded)

Recommended Queries

Judge vs Human Agreement Benchmark Slice: DROP Metric Slice: accuracy IAA-Reported Evaluations Recent High-Signal Papers

Start with These 3

Use these when you need one protocol anchor, one benchmark anchor, and one recent comparison point before reading the wider hub.

Strongest protocol reference

Personalized RewardBench: Evaluating Reward Models with Human Aligned…

Highest protocol score with explicit human/eval signal plus Rewardbench.

Strongest benchmark reference

PubMed Reasoner: Dynamic Reasoning-based Retrieval for Evidence-Groun…

MMLU with accuracy gives a fast comparison anchor.

Strongest recent paper

TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step To…

Useful for current practice scanning; published Apr 8, 2026.

Start Here (Best First 6)

Ranked for protocol completeness (human signal, benchmark + metric anchors, quality controls, and judge/human overlap).

Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization
Apr 8, 2026 · Citations: 0 · Score: 7.5

HF: Pairwise Preference, Rubric Rating · Eval: Human Eval, Automatic Metrics · Benchmark: Rewardbench · Metric: Accuracy
PubMed Reasoner: Dynamic Reasoning-based Retrieval for Evidence-Grounded Biomedical Question Answering
Mar 28, 2026 · Citations: 0 · Score: 7.5

HF: Expert Verification · Eval: Llm As Judge, Automatic Metrics · Benchmark: MMLU · Metric: Accuracy
TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories
Apr 8, 2026 · Citations: 0 · Score: 7.5

HF: Red Team · Eval: Automatic Metrics · Benchmark: Tracesafe Bench · Metric: Accuracy
Beyond Paper-to-Paper: Structured Profiling and Rubric Scoring for Paper-Reviewer Matching
Apr 7, 2026 · Citations: 0 · Score: 7.5

HF: Rubric Rating · Eval: Automatic Metrics · Benchmark: Scirepeval · Metric: Recall
Paper Reconstruction Evaluation: Evaluating Presentation and Hallucination in AI-written Papers
Apr 1, 2026 · Citations: 0 · Score: 7.5

HF: Rubric Rating · Eval: Automatic Metrics · Benchmark: Paperwrite Bench · Metric: Cost
Rethinking Atomic Decomposition for LLM Judges: A Prompt-Controlled Study of Reference-Grounded QA Evaluation
Mar 30, 2026 · Citations: 0 · Score: 7.5

HF: Rubric Rating · Eval: Automatic Metrics · Benchmark: TruthfulQA · Metric: Accuracy

Protocol Matrix (Top 12)

Use this to quickly compare protocol ingredients instead of scanning long prose.

Paper	HF Signal	Eval Modes	Benchmarks	Metrics	QC
Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization Apr 8, 2026	Yes Pairwise Preference , Rubric Rating	Human Eval , Automatic Metrics	Rewardbench	Accuracy , Helpfulness	Not Reported
PubMed Reasoner: Dynamic Reasoning-based Retrieval for Evidence-Grounded Biomedical Question Answering Mar 28, 2026	Yes Expert Verification	Llm As Judge , Automatic Metrics	MMLU	Accuracy , Relevance	Not Reported
TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories Apr 8, 2026	Yes Red Team	Automatic Metrics	Tracesafe Bench	Accuracy	Not Reported
Beyond Paper-to-Paper: Structured Profiling and Rubric Scoring for Paper-Reviewer Matching Apr 7, 2026	Yes Rubric Rating	Automatic Metrics	Scirepeval	Recall	Not Reported
Paper Reconstruction Evaluation: Evaluating Presentation and Hallucination in AI-written Papers Apr 1, 2026	Yes Rubric Rating	Automatic Metrics	Paperwrite Bench	Not Reported	Not Reported
Rethinking Atomic Decomposition for LLM Judges: A Prompt-Controlled Study of Reference-Grounded QA Evaluation Mar 30, 2026	Yes Rubric Rating	Automatic Metrics	TruthfulQA	Accuracy	Not Reported
Do Phone-Use Agents Respect Your Privacy? Apr 1, 2026	Yes Pairwise Preference	Automatic Metrics	APPS , Myphonebench	Task success	Not Reported
ReDAct: Uncertainty-Aware Deferral for LLM Agents Apr 8, 2026	No Not Reported	Simulation Env	ALFWorld	Token cost	Not Reported
DataSTORM: Deep Research on Large-Scale Databases using Exploratory Data Analysis and Data Storytelling Apr 7, 2026	No Not Reported	Human Eval	Insightbench	Recall	Not Reported
LUDOBENCH: Evaluating LLM Behavioural Decision-Making Through Spot-Based Board Game Scenarios in Ludo Apr 7, 2026	No Not Reported	Simulation Env	Ludobench	Dice	Not Reported
LLM-as-a-Judge for Time Series Explanations Apr 2, 2026	No Not Reported	Llm As Judge , Automatic Metrics	DROP	Accuracy , Faithfulness	Not Reported
Navigating Large-Scale Document Collections: MuDABench for Multi-Document Analytical QA Apr 24, 2026	No Not Reported	Automatic Metrics	Mudabench	Accuracy	Not Reported

Protocol Diff (Top Papers)

Fast side-by-side comparison for the highest-ranked papers in this hub.

Signal	Personalized RewardBench: Evaluating Reward Models…	PubMed Reasoner: Dynamic Reasoning-based Retrieval…	TraceSafe: A Systematic Assessment of LLM Guardrail…
Human Feedback	Pairwise Preference, Rubric Rating	Expert Verification	Red Team
Evaluation Modes	Human Eval, Automatic Metrics	Llm As Judge, Automatic Metrics	Automatic Metrics
Benchmarks	Rewardbench	MMLU	Tracesafe Bench
Metrics	Accuracy, Helpfulness	Accuracy, Relevance	Accuracy
Quality Controls	Not reported	Not reported	Not reported
Rater Population	Unknown	Domain Experts	Unknown
Annotation Unit	Pairwise	Unknown	Trajectory

Research Utility Snapshot

Human Feedback Mix

Pairwise Preference (39)
Expert Verification (21)
Rubric Rating (15)
Critique Edit (11)

Evaluation Modes

Automatic Metrics (271)
Simulation Env (17)
Human Eval (16)
Llm As Judge (16)

Top Benchmarks

DROP (15)
MMLU (12)
GSM8K (9)
SemEval (7)

Top Metrics

Accuracy (181)
Cost (81)
F1 (40)
Agreement (39)

Rater Population Mix

Domain Experts (81)
Mixed (5)

Quality Controls

Calibration (24)
Inter Annotator Agreement Reported (11)
Adjudication (7)
Gold Questions (5)

Coverage diagnostics (sample-based): human-feedback 81.7% · benchmarks 36.7% · metrics 75.0% · quality controls 11.7%.

Top Papers

Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization
Qiyao Ma, Dechen Gao, Rui Cai, Boqi Zhao, Hanchu Zhou · Apr 8, 2026 · Citations: 0

Pairwise PreferenceRubric Rating Human EvalAutomatic Metrics

Pluralistic alignment has emerged as a critical frontier in the development of Large Language Models (LLMs), with reward models (RMs) serving as a central mechanism for capturing diverse human values.
PubMed Reasoner: Dynamic Reasoning-based Retrieval for Evidence-Grounded Biomedical Question Answering
Yiqing Zhang, Xiaozhong Liu, Fabricio Murai · Mar 28, 2026 · Citations: 0

Expert Verification Llm As JudgeAutomatic Metrics

In this context, we introduce PubMed Reasoner, a biomedical QA agent composed of three stages: self-critic query refinement evaluates MeSH terms for coverage, alignment, and redundancy to enhance PubMed queries based on partial (metadata)…
TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories
Yen-Shan Chen, Sian-Yao Huang, Cheng-Lin Yang, Yun-Nung Chen · Apr 8, 2026 · Citations: 0

Red Team Automatic Metrics Long Horizon

As large language models (LLMs) evolve from static chatbots into autonomous agents, the primary vulnerability surface shifts from final outputs to intermediate execution traces.
Beyond Paper-to-Paper: Structured Profiling and Rubric Scoring for Paper-Reviewer Matching
Yicheng Pan, Zhiyuan Ning, Ludi Wang, Yi Du · Apr 7, 2026 · Citations: 0

Rubric Rating Automatic Metrics

To address this gap, we propose P2R, a training-free framework that shifts from implicit paper-to-paper matching to explicit profile-based matching.
Paper Reconstruction Evaluation: Evaluating Presentation and Hallucination in AI-written Papers
Atsuyuki Miyai, Mashiro Toyooka, Zaiying Zhao, Kenta Watanabe, Toshihiko Yamasaki · Apr 1, 2026 · Citations: 0

Rubric Rating Automatic Metrics

We introduce Paper Reconstruction Evaluation (PaperRecon), an evaluation framework in which an overview (overview.md) is created from an existing paper, after which an agent generates a full paper based on the overview and minimal…
Rethinking Atomic Decomposition for LLM Judges: A Prompt-Controlled Study of Reference-Grounded QA Evaluation
Xinran Zhang · Mar 30, 2026 · Citations: 0

Rubric Rating Automatic Metrics

Atomic decomposition -- breaking a candidate answer into claims before verifying each against a reference -- is a widely adopted design for LLM-based reference-grounded judges.
ReDAct: Uncertainty-Aware Deferral for LLM Agents
Dzianis Piatrashyn, Nikita Kotelevskii, Kirill Grishchenkov, Nikita Glazkov, Ivan Nasonov · Apr 8, 2026 · Citations: 0

Simulation Env Long Horizon

Recently, LLM-based agents have become increasingly popular across many applications, including complex sequential decision-making problems.
Do Phone-Use Agents Respect Your Privacy?
Zhengyang Tang, Ke Ji, Xidong Wang, Zihan Ye, Xinyuan Wang · Apr 1, 2026 · Citations: 0

Pairwise Preference Automatic Metrics

We study whether phone-use agents respect privacy while completing benign mobile tasks.
DataSTORM: Deep Research on Large-Scale Databases using Exploratory Data Analysis and Data Storytelling
Shicheng Liu, Yucheng Jiang, Sajid Farook, Camila Nicollier Sanchez, David Fernando Castro Pena · Apr 7, 2026 · Citations: 0

Human Eval Long Horizon

Deep research with Large Language Model (LLM) agents is emerging as a powerful paradigm for multi-step information discovery, synthesis, and analysis.
LUDOBENCH: Evaluating LLM Behavioural Decision-Making Through Spot-Based Board Game Scenarios in Ludo
Ojas Jain, Dhruv Kumar · Apr 7, 2026 · Citations: 0

Simulation Env Multi Agent

We introduce LudoBench, a benchmark for evaluating LLM strategic reasoning in Ludo, a stochastic multi-agent board game whose dice mechanics, piece capture, safe-square navigation, and home-path progression introduce meaningful planning…
LLM-as-a-Judge for Time Series Explanations
Preetham Sivalingam, Murari Mandal, Saurabh Deshpande, Dhruv Kumar · Apr 2, 2026 · Citations: 0

Llm As JudgeAutomatic Metrics

Although modern models generate textual interpretations of numerical signals, existing evaluation methods are limited: reference based similarity metrics and consistency checking models require ground truth explanations, while traditional…
Navigating Large-Scale Document Collections: MuDABench for Multi-Document Analytical QA
Zhanli Li, Yixuan Cao, Lvzhou Luo, Ping Luo · Apr 24, 2026 · Citations: 0

Automatic Metrics Multi Agent

We present MuDABench, a benchmark for multi-document analytical QA, where questions require extracting and synthesizing information across numerous documents to perform quantitative analysis.
Don't Overthink It: Inter-Rollout Action Agreement as a Free Adaptive-Compute Signal for LLM Agents
Khushal Sethi · Apr 9, 2026 · Citations: 0

Automatic Metrics Long Horizon

We introduce TrACE (Trajectorical Adaptive Compute via agrEement), a training-free controller that allocates LLM calls adaptively across agent timesteps by measuring inter-rollout action agreement.
Brief Is Better: Non-Monotonic Chain-of-Thought Budget Effects in Function-Calling Language Agents
Xuan Qi · Apr 2, 2026 · Citations: 0

Automatic Metrics Tool Use

Chain-of-thought (CoT) reasoning is widely assumed to improve agent performance, but the relationship between reasoning length and accuracy in structured tool-use settings remains poorly understood.
OSCAR: Orchestrated Self-verification and Cross-path Refinement
Yash Shah, Abhijit Chakraborty, Naresh Kumar Devulapally, Vishnu Lokhande, Vivek Gupta · Apr 2, 2026 · Citations: 0

Automatic Metrics Long Horizon

We introduce a suite of trajectory-level assessments, including a cross-chain divergence-at-hallucination (CDH) metric, for principled comparison of localization methods.
S0 Tuning: Zero-Overhead Adaptation of Hybrid Recurrent-Attention Models
Jack Young · Apr 1, 2026 · Citations: 0

Automatic Metrics Long Horizon

Using roughly 48 execution-verified HumanEval training solutions, tuning a single initial state matrix per recurrent layer, with zero inference overhead, outperforms LoRA by +10.8 pp (p < 0.001) on HumanEval.
Asymmetric Actor-Critic for Multi-turn LLM Agents
Shuli Jiang, Zhaoyang Zhang, Yi Zhang, Shuo Yang, Wei Xia · Mar 31, 2026 · Citations: 0

Automatic Metrics Long Horizon

In many real-world applications, agents must succeed in one-shot settings where retries are impossible.

Related Hubs

Get Started

Join the #1 Platform for AI Training Talent

Where top AI builders and expert AI Trainers connect to build the future of AI.

Self-Service

Post a Job

Post your project and get a shortlist of qualified AI Trainers and Data Labelers. Hire and manage your team in the tools you already use.

Create Account & Post a Job

Managed Service

For Large Projects

Done-for-You

We recruit, onboard, and manage a dedicated team inside your tools. End-to-end operations for large or complex projects.

Learn About Managed Service

For Freelancers

Join as an AI Trainer

Find AI training and data labeling projects across platforms, all in one place. One profile, one application process, more opportunities.

Join Now