HFEPX Benchmark Hub

Reasoning & Math Suite Benchmark Papers + Automatic Metrics

Updated from current HFEPX corpus (Mar 17, 2026). 10 papers are grouped in this benchmark page.

Read Full Context

Updated from current HFEPX corpus (Mar 17, 2026). 10 papers are grouped in this benchmark page. Common evaluation modes: Automatic Metrics. Common annotation unit: Trajectory. Frequently cited benchmark: GSM8K. Common metric signal: accuracy. Use this page to compare protocol setup, judge behavior, and labeling design decisions before running new eval experiments. Newest paper in this set is from Mar 4, 2026.

Papers: 10 Last published: Mar 4, 2026 Global RSS

Researcher Quick Triage

Use this page for benchmark-matched method comparisons and eval protocol selection. Quality band: Developing .

High-Signal Coverage

100.0%

10 / 10 sampled papers are not low-signal flagged.

Replication-Ready Set

Papers with explicit benchmark + metric + eval mode fields.

Quality Controls

0.0%

0 papers report calibration/adjudication/IAA controls.

10 papers explicitly name benchmark datasets in the sampled set.
10 papers report at least one metric term in metadata extraction.
Start with the ranked shortlist below before reading all papers.

Primary action: Start with the top 2 benchmark-matched papers, then compare evaluation modes in the protocol matrix.

Why This Matters (Expanded)

Why This Matters For Eval Research

40% of papers report explicit human-feedback signals, led by pairwise preferences.
automatic metrics appears in 100% of papers in this hub.
GSM8K is a recurring benchmark anchor for cross-paper comparisons in this page.

Protocol Notes (Expanded)

Protocol Takeaways

Quality-control reporting is sparse in this slice; prioritize papers with explicit calibration or adjudication steps.
Rater context is mostly unspecified rater pools, and annotation is commonly trajectory-level annotation; use this to scope replication staffing.
Stratify by benchmark (GSM8K vs MMLU) before comparing methods.

Benchmark Interpretation

GSM8K appears in 40% of hub papers (4/10); use this cohort for benchmark-matched comparisons.
MMLU appears in 40% of hub papers (4/10); use this cohort for benchmark-matched comparisons.

Metric Interpretation

accuracy is reported in 70% of hub papers (7/10); compare with a secondary metric before ranking methods.
cost is reported in 40% of hub papers (4/10); compare with a secondary metric before ranking methods.

Start Here (Benchmark-Matched First 6)

Ranked by protocol completeness so you can quickly find papers suitable for comparison studies.

$V_1$: Unifying Generation and Self-Verification for Parallel Reasoners
Mar 4, 2026 · Citations: 0 · Score: 8.5

Eval: Automatic Metrics · Metrics: Pass@1
How Reliable is Language Model Micro-Benchmarking?
Oct 9, 2025 · Citations: 0 · Score: 7.5

Eval: Automatic Metrics · Metrics: Accuracy
FOR-Prompting: From Objection to Revision via an Asymmetric Prompting Protocol
Oct 2, 2025 · Citations: 0 · Score: 7.5

Eval: Automatic Metrics · Metrics: Accuracy
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Jun 3, 2025 · Citations: 0 · Score: 7.0

Eval: Automatic Metrics · Metrics: Pass@1
Top-b: Entropic Regulation of Relative Probability Bands in Autoregressive Language Processes
Mar 15, 2026 · Citations: 0 · Score: 7.0

Eval: Automatic Metrics · Metrics: Accuracy
Learning When to Sample: Confidence-Aware Self-Consistency for Efficient LLM Chain-of-Thought Reasoning
Mar 9, 2026 · Citations: 0 · Score: 7.0

Eval: Automatic Metrics · Metrics: Accuracy

Protocol Matrix (Top 10)

Compare protocol ingredients quickly before deep-reading full papers.

Paper	Eval Modes	Human Feedback	Metrics	Quality Controls
$V_1$: Unifying Generation and Self-Verification for Parallel Reasoners Mar 4, 2026	Automatic Metrics	Pairwise Preference	Pass@1	Not reported
How Reliable is Language Model Micro-Benchmarking? Oct 9, 2025	Automatic Metrics	Pairwise Preference	Accuracy, Cost	Not reported
FOR-Prompting: From Objection to Revision via an Asymmetric Prompting Protocol Oct 2, 2025	Automatic Metrics	Pairwise Preference, Critique Edit	Accuracy	Not reported
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Jun 3, 2025	Automatic Metrics	Critique Edit	Pass@1	Not reported
Top-b: Entropic Regulation of Relative Probability Bands in Autoregressive Language Processes Mar 15, 2026	Automatic Metrics	Not reported	Accuracy	Not reported
Learning When to Sample: Confidence-Aware Self-Consistency for Efficient LLM Chain-of-Thought Reasoning Mar 9, 2026	Automatic Metrics	Not reported	Accuracy, Cost	Not reported
D-COT: Disciplined Chain-of-Thought Learning for Efficient Reasoning in Small Language Models Feb 25, 2026	Automatic Metrics	Not reported	Accuracy	Not reported
Confidence-Driven Multi-Scale Model Selection for Cost-Efficient Inference Feb 25, 2026	Automatic Metrics	Not reported	Accuracy, Cost	Not reported
Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs Dec 3, 2025	Automatic Metrics	Not reported	Cost	Not reported
SPARE: Single-Pass Annotation with Reference-Guided Evaluation for Automatic Process Supervision and Reward Modelling Jun 18, 2025	Automatic Metrics	Not reported	Accuracy, Precision	Not reported

Researcher Workflow (Detailed)

Checklist

Moderate: Papers with explicit human feedback

Coverage is usable but incomplete (40% vs 45% target).
Gap: Papers reporting quality controls

Coverage is a replication risk (0% vs 30% target).
Strong: Papers naming benchmarks/datasets

Coverage is strong (100% vs 35% target).
Strong: Papers naming evaluation metrics

Coverage is strong (100% vs 35% target).
Gap: Papers with known rater population

Coverage is a replication risk (0% vs 35% target).
Strong: Papers with known annotation unit

Coverage is strong (70% vs 35% target).

Strengths

Most papers provide measurable evaluation context (100% benchmarks, 100% metrics).
Agentic evaluation appears in 60% of papers.

Known Gaps

Only 0% of papers report quality controls; prioritize calibration/adjudication evidence.
Rater population is under-specified (0% coverage).

Suggested Next Analyses

Stratify by benchmark (GSM8K vs MMLU) before comparing methods.
Track metric sensitivity by reporting both accuracy and cost.

Recommended Queries

Benchmark Slice: GSM8K Metric Slice: accuracy Recent High-Signal Papers

Known Limitations

Only 0% of papers report quality controls; prioritize calibration/adjudication evidence.
Rater population is under-specified (0% coverage).
Narrative synthesis is grounded in metadata and abstracts only; full-paper implementation details are not parsed.

Research Utility Snapshot (Detailed)

Evaluation Modes

Automatic Metrics (10)

Human Feedback Mix

Pairwise Preference (3)
Critique Edit (2)

Top Benchmarks

GSM8K (4)
MMLU (4)
AIME (2)
GPQA (2)

Top Metrics

Accuracy (7)
Cost (4)
Pass@1 (2)
Inference cost (1)

Top Papers On This Benchmark

$V_1$: Unifying Generation and Self-Verification for Parallel Reasoners
Harman Singh, Xiuyu Li, Kusha Sareen, Monishwaran Maheswaran, Sijun Tan · Mar 4, 2026 · Citations: 0

Pairwise Preference Automatic Metrics

On code generation (LiveCodeBench, CodeContests, SWE-Bench) and math reasoning (AIME, HMMT) benchmarks, V_1-Infer improves Pass@1 by up to 10% over pointwise verification and outperforms recent test-time scaling methods while being…
How Reliable is Language Model Micro-Benchmarking?
Gregory Yauney, Shahzaib Saqib Warraich, Swabha Swayamdipta · Oct 9, 2025 · Citations: 0

Pairwise Preference Automatic Metrics

We introduce a meta-evaluation measure for micro-benchmarking which investigates how well a micro-benchmark can rank two models as a function of their performance difference on the full benchmark.
FOR-Prompting: From Objection to Revision via an Asymmetric Prompting Protocol
He Zhang, Anzhou Zhang, Jian Dai · Oct 2, 2025 · Citations: 0

Pairwise PreferenceCritique Edit Automatic Metrics

Beyond structured math tasks, FOR-Prompting supports refinement in open-ended and multi-stage tasks: qualitative analysis shows improved exploration, coverage, and specificity, and a blind study of human preferences found that participants…
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Xiaoying Zhang, Yipeng Zhang, Hao Sun, Kaituo Feng, Chaochao Lu · Jun 3, 2025 · Citations: 0

Critique Edit Automatic Metrics

We show that plateaued RL models can successfully refine failed solutions when given natural language critiques.
Top-b: Entropic Regulation of Relative Probability Bands in Autoregressive Language Processes
Deepon Halder, Raj Dabre · Mar 15, 2026 · Citations: 0

Automatic Metrics

Empirical validation on GPQA and GSM8K benchmarks indicates that Top-b significantly reduces generation entropy and inter-decoding variance while maintaining competitive reasoning accuracy, effectively approximating a self-regulating…
Learning When to Sample: Confidence-Aware Self-Consistency for Efficient LLM Chain-of-Thought Reasoning
Juming Xiong, Kevin Guo, Congning Ni, Chao Yan, Katherine Brown · Mar 9, 2026 · Citations: 0

Automatic Metrics

Recent self-consistency-based approaches further improve accuracy but require sampling and aggregating multiple reasoning trajectories, leading to substantial additional computational overhead.
D-COT: Disciplined Chain-of-Thought Learning for Efficient Reasoning in Small Language Models
Shunsuke Ubukata · Feb 25, 2026 · Citations: 0

Automatic Metrics

In this study, we propose Disciplined Chain-of-Thought (D-CoT), a novel framework that enforces a structured reasoning process using control tags -- such as <TEMP_LOW> for fact-checking and <TEMP_HIGH> for multi-perspective exploration --…
Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs
Ngoc Bui, Shubham Sharma, Simran Lamba, Saumitra Mishra, Rex Ying · Dec 3, 2025 · Citations: 0

Automatic Metrics

Across mathematical reasoning (GSM8K, MATH-500, AIME24), procedural generation (LongProc), conversational long-memory benchmarks (LongMemEval), and long-context understanding (LongBenchV2 and SCBench), TRIM-KV consistently outperforms…
SPARE: Single-Pass Annotation with Reference-Guided Evaluation for Automatic Process Supervision and Reward Modelling
Md Imbesat Hassan Rizvi, Xiaodan Zhu, Iryna Gurevych · Jun 18, 2025 · Citations: 0

Automatic Metrics

To address this, we introduce Single-Pass Annotation with Reference-Guided Evaluation (SPARE), a novel structured framework that enables efficient per-step annotation by jointly aligning solution steps to reference solutions and determine…
Confidence-Driven Multi-Scale Model Selection for Cost-Efficient Inference
Bo-Wei Chen, Chung-Chi Chen, An-Zi Yen · Feb 25, 2026 · Citations: 0

Automatic Metrics

Experiments on the Massive Multitask Language Understanding (MMLU) benchmark show that our approach achieves accuracy comparable to the largest model while reducing computational costs by 20\% to 40\%.

Related Benchmark Hubs

Reasoning & Math Suite Benchmark Papers Reasoning & Math Suite Benchmark Papers In CS.CL General Or Math Papers Automatic Metrics Papers Automatic Metrics Or Human Eval Papers Automatic Metrics Or Llm As Judge Papers DROP Benchmark Papers (Last 45 Days) (11) DROP Benchmark Papers (Last 60 Days) (11) DROP Benchmark Papers (Last 75 Days) (11) Reasoning & Math Suite Benchmark Papers (23) Reasoning & Math Suite Benchmark Papers In CS.CL (21) Reasoning & Math Suite Benchmark Papers In CS.AI (15) Coding Evaluation Suite Benchmark Papers (11) Coding Evaluation Suite Benchmark Papers In CS.CL (10) SWE-Bench Ecosystem Benchmark Papers (13) SWE-Bench Ecosystem Benchmark Papers In CS.CL (11)

Need human evaluators for your AI research? Scale annotation with expert AI Trainers.

Post a Job Get a Quote