HFEPX Benchmark Hub

MATH-500 Or GPQA Benchmark Papers

Updated from current HFEPX corpus (Jun 30, 2026). 45 papers are grouped in this benchmark page.

Read Full Context

Updated from current HFEPX corpus (Jun 30, 2026). 45 papers are grouped in this benchmark page. Common evaluation modes: Automatic Metrics, Llm As Judge. Common annotation unit: Trajectory. Frequently cited benchmark: MATH-500. Common metric signal: accuracy. Use this page to compare protocol setup, judge behavior, and labeling design decisions before running new eval experiments. Newest paper in this set is from Jun 24, 2026.

Papers: 45 Last published: Jun 24, 2026 Global RSS

Researcher Quick Triage

Use this page for benchmark-matched method comparisons and eval protocol selection. Quality band: High .

High-Signal Coverage

100.0%

45 / 45 sampled papers are not low-signal flagged.

Replication-Ready Set

Papers with explicit benchmark + metric + eval mode fields.

Quality Controls

0.0%

0 papers report calibration/adjudication/IAA controls.

12 papers explicitly name benchmark datasets in the sampled set.
11 papers report at least one metric term in metadata extraction.
Start with the ranked shortlist below before reading all papers.

Primary action: Start with the top 2 benchmark-matched papers, then compare evaluation modes in the protocol matrix.

Why This Matters (Expanded)

Why This Matters For Eval Research

6.7% of papers report explicit human-feedback signals, led by pairwise preferences.
automatic metrics appears in 24.4% of papers in this hub.
MATH-500 is a recurring benchmark anchor for cross-paper comparisons in this page.

Protocol Notes (Expanded)

Protocol Takeaways

Quality-control reporting is sparse in this slice; prioritize papers with explicit calibration or adjudication steps.
Rater context is mostly unspecified rater pools, and annotation is commonly trajectory-level annotation; use this to scope replication staffing.
Pair this hub with a human_eval-heavy hub to validate judge-model calibration.

Benchmark Interpretation

MATH-500 appears in 57.8% of hub papers (26/45); use this cohort for benchmark-matched comparisons.
GPQA appears in 53.3% of hub papers (24/45); use this cohort for benchmark-matched comparisons.

Metric Interpretation

accuracy is reported in 44.4% of hub papers (20/45); compare with a secondary metric before ranking methods.
cost is reported in 20% of hub papers (9/45); compare with a secondary metric before ranking methods.

Start Here (Benchmark-Matched First 6)

Ranked by protocol completeness so you can quickly find papers suitable for comparison studies.

Cliff Tokens: Identifying Single-Token Failure Triggers in LLM Mathematical Reasoning
Jun 24, 2026 · Citations: 0 · Score: 8.5

Eval: Automatic Metrics · Metrics: Accuracy
Learning How to Use Tools, Not Just When: Pattern-Aware Tool-Integrated Reasoning
Sep 27, 2025 · Citations: 0 · Score: 7.0

Eval: Automatic Metrics · Metrics: Accuracy
Blockwise Policy-Drift Gating for On-Policy Distillation
Jun 23, 2026 · Citations: 0 · Score: 7.0

Eval: Automatic Metrics · Metrics: Pass@8
S0 Tuning: Zero-Overhead Adaptation of Hybrid Recurrent-Attention Models
Apr 1, 2026 · Citations: 0 · Score: 6.0

Eval: Automatic Metrics · Metrics: Pass@1
Top-b: Entropic Regulation of Relative Probability Bands in Autoregressive Language Processes
Mar 15, 2026 · Citations: 0 · Score: 6.0

Eval: Automatic Metrics · Metrics: Accuracy
D-COT: Disciplined Chain-of-Thought Learning for Efficient Reasoning in Small Language Models
Feb 25, 2026 · Citations: 0 · Score: 6.0

Eval: Automatic Metrics · Metrics: Accuracy

Protocol Matrix (Top 10)

Compare protocol ingredients quickly before deep-reading full papers.

Paper	Eval Modes	Human Feedback	Metrics	Quality Controls
Cliff Tokens: Identifying Single-Token Failure Triggers in LLM Mathematical Reasoning Jun 24, 2026	Automatic Metrics	Pairwise Preference	Accuracy, Pass@64	Not reported
Learning How to Use Tools, Not Just When: Pattern-Aware Tool-Integrated Reasoning Sep 27, 2025	Automatic Metrics	Pairwise Preference	Accuracy	Not reported
Blockwise Policy-Drift Gating for On-Policy Distillation Jun 23, 2026	Automatic Metrics	Not reported	Pass@8	Not reported
S0 Tuning: Zero-Overhead Adaptation of Hybrid Recurrent-Attention Models Apr 1, 2026	Automatic Metrics	Not reported	Pass@1, Cost	Not reported
Top-b: Entropic Regulation of Relative Probability Bands in Autoregressive Language Processes Mar 15, 2026	Automatic Metrics	Not reported	Accuracy	Not reported
D-COT: Disciplined Chain-of-Thought Learning for Efficient Reasoning in Small Language Models Feb 25, 2026	Automatic Metrics	Not reported	Accuracy	Not reported
Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs Dec 3, 2025	Automatic Metrics	Not reported	Cost	Not reported
DeepPrune: Parallel Scaling without Inter-trace Redundancy Oct 9, 2025	Llm As Judge, Automatic Metrics	Not reported	Accuracy, Auroc	Not reported
SIGMA: Search-Augmented On-Demand Knowledge Integration for Agentic Mathematical Reasoning Oct 31, 2025	Automatic Metrics	Not reported	Accuracy	Not reported
Schema for In-Context Learning Oct 14, 2025	Not reported	Demonstrations	Not reported	Not reported

Researcher Workflow (Detailed)

Checklist

Gap: Papers with explicit human feedback

Coverage is a replication risk (6.7% vs 45% target).
Gap: Papers reporting quality controls

Coverage is a replication risk (0% vs 30% target).
Strong: Papers naming benchmarks/datasets

Coverage is strong (100% vs 35% target).
Strong: Papers naming evaluation metrics

Coverage is strong (75.6% vs 35% target).
Gap: Papers with known rater population

Coverage is a replication risk (0% vs 35% target).
Gap: Papers with known annotation unit

Coverage is a replication risk (13.3% vs 35% target).

Strengths

Most papers provide measurable evaluation context (100% benchmarks, 75.6% metrics).

Known Gaps

Only 0% of papers report quality controls; prioritize calibration/adjudication evidence.
Rater population is under-specified (0% coverage).
Annotation unit is under-specified (13.3% coverage).

Suggested Next Analyses

Pair this hub with a human_eval-heavy hub to validate judge-model calibration.
Stratify by benchmark (MATH-500 vs GPQA) before comparing methods.
Track metric sensitivity by reporting both accuracy and cost.

Recommended Queries

LLM-as-Judge Protocols Benchmark Slice: MATH-500 Metric Slice: accuracy Recent High-Signal Papers

Known Limitations

Only 0% of papers report quality controls; prioritize calibration/adjudication evidence.
Rater population is under-specified (0% coverage).
Narrative synthesis is grounded in metadata and abstracts only; full-paper implementation details are not parsed.

Research Utility Snapshot (Detailed)

Evaluation Modes

Automatic Metrics (11)
Llm As Judge (1)

Human Feedback Mix

Pairwise Preference (2)
Demonstrations (1)

Top Benchmarks

MATH 500 (26)
GPQA (24)
AIME (16)
GSM8K (11)

Top Metrics

Accuracy (20)
Cost (9)
Latency (3)
Pass@1 (3)

Top Papers On This Benchmark

Cliff Tokens: Identifying Single-Token Failure Triggers in LLM Mathematical Reasoning
Jaeyong Ko, Pilsung Kang, Yukyung Lee · Jun 24, 2026 · Citations: 0

Pairwise Preference Automatic Metrics

Across seven models and three mathematical reasoning benchmarks (GSM1K, MATH500, AIME 2025), cliff tokens act as failure triggers; deleting the first cliff token and resampling recovers pass@64 to 1.0, while keeping it limits recovery to…
Learning How to Use Tools, Not Just When: Pattern-Aware Tool-Integrated Reasoning
Ningning Xu, Yuxuan Jiang, Shubhashis Roy Dipta, Hengyuan Zhang · Sep 27, 2025 · Citations: 0

Pairwise Preference Automatic Metrics

We propose a two-stage framework that first builds code competence from both patterns and then aligns pattern selection with teacher preferences.
Blockwise Policy-Drift Gating for On-Policy Distillation
Liwen Zheng, Haiyun Jiang · Jun 23, 2026 · Citations: 0

Automatic Metrics

In a six-variant Qwen3 math reasoning benchmark with a uniform 200-step training budget for all trained variants, we use pass@8 as the primary problem-level solve-rate metric.
S0 Tuning: Zero-Overhead Adaptation of Hybrid Recurrent-Attention Models
Jack Young · Apr 1, 2026 · Citations: 0

Automatic Metrics

Using roughly 48 execution-verified HumanEval training solutions, tuning a single initial state matrix per recurrent layer, with zero inference overhead, outperforms LoRA by +10.8 pp (p < 0.001) on HumanEval.
Top-b: Entropic Regulation of Relative Probability Bands in Autoregressive Language Processes
Deepon Halder, Raj Dabre · Mar 15, 2026 · Citations: 0

Automatic Metrics

Empirical validation on GPQA and GSM8K benchmarks indicates that Top-b significantly reduces generation entropy and inter-decoding variance while maintaining competitive reasoning accuracy, effectively approximating a self-regulating…
D-COT: Disciplined Chain-of-Thought Learning for Efficient Reasoning in Small Language Models
Shunsuke Ubukata · Feb 25, 2026 · Citations: 0

Automatic Metrics

In this study, we propose Disciplined Chain-of-Thought (D-CoT), a novel framework that enforces a structured reasoning process using control tags -- such as <TEMP_LOW> for fact-checking and <TEMP_HIGH> for multi-perspective exploration --…
Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs
Ngoc Bui, Shubham Sharma, Simran Lamba, Saumitra Mishra, Rex Ying · Dec 3, 2025 · Citations: 0

Automatic Metrics

Across mathematical reasoning (GSM8K, MATH-500, AIME24), procedural generation (LongProc), conversational long-memory benchmarks (LongMemEval), and long-context understanding (LongBenchV2 and SCBench), TRIM-KV consistently outperforms…
Accelerated Test-Time Scaling with Model-Free Speculative Sampling
Woomin Song, Saket Dingliwal, Sai Muralidhar Jayanthi, Bhavana Ganesh, Jinwoo Shin · Jun 5, 2025 · Citations: 0

Automatic Metrics

Extensive evaluations across multiple models and reasoning tasks (AIME-2024, GPQA-Diamond, and LiveCodeBench) demonstrate that STAND reduces inference latency by 60-65% compared to standard autoregressive decoding while maintaining…
DeepPrune: Parallel Scaling without Inter-trace Redundancy
Shangqing Tu, Yaxuan Li, Yushi Bai, Lei Hou, Juanzi Li · Oct 9, 2025 · Citations: 0

Llm As JudgeAutomatic Metrics

Our method features a specialized judge model trained with out-of-distribution data (AIME 2022, AIME 2023, and MATH 500) using oversampling techniques to accurately predict answer equivalence from partial reasoning traces, achieving 0.7072…
SIGMA: Search-Augmented On-Demand Knowledge Integration for Agentic Mathematical Reasoning
Ali Asgarov, Umid Suleymanov, Aadyant Khatri · Oct 31, 2025 · Citations: 0

Automatic Metrics

We introduce SIGMA (Search-Augmented On-Demand Knowledge Integration for AGentic Mathematical reAsoning), a unified framework that orchestrates specialized agents to independently reason, perform targeted searches, and synthesize findings…
Towards Hierarchical Multi-Step Reward Models for Enhanced Reasoning in Large Language Models
Teng Wang, Zhangyi Jiang, Zhenqi He, Shenyang Tong, Wenhan Yang · Mar 16, 2025 · Citations: 0

Automatic Metrics

Empirical results on the PRM800K dataset show that HRM, together with HNC, provides more stable and reliable evaluations than PRM.
Schema for In-Context Learning
Pan Chen, Shaohong Chen, Mark Wang, Shi Xuan Leong, Priscilla Fung · Oct 14, 2025 · Citations: 0

Demonstrations

Inspired by cognitive science, specifically schema theory, which holds that humans interpret new information by activating pre-existing mental frameworks (schemas) to structure understanding, we introduce Schema-Activated In-Context…
ACC: Compiling Agent Trajectories for Long-Context Training
Qisheng Su, Zhen Fang, Shiting Huang, Yu Zeng, Yiming Zhao · May 21, 2026 · Citations: 0
LamPO: A Lambda Style Policy Optimization for Reasoning Language Models
Zhe Yuan, Yipeng Zhou, Jinghan Li, Xinyuan Chen, Bowen Deng · May 20, 2026 · Citations: 0
Weak-to-Strong Elicitation via Mismatched Wrong Drafts
Wei Deng · May 17, 2026 · Citations: 0
Minimal-Intervention KV Retention via Set-Conditioned Diversity
Libo Sun, Po-wei Harn, Peixiong He, Xiao Qin · May 14, 2026 · Citations: 0
FocuSFT: Bilevel Optimization for Dilution-Aware Long-Context Fine-Tuning
Zehua Pei, Hui-Ling Zhen, Xianzhi Yu, Sinno Jialin Pan, Mingxuan Yuan · May 11, 2026 · Citations: 0
The Metacognitive Probe: Five Behavioural Calibration Diagnostics for LLMs
Rafael C. T. Oliveira · May 11, 2026 · Citations: 0
Not All Thoughts Need HBM: Semantics-Aware Memory Hierarchy for LLM Reasoning
Aojie Yuan, Tianqi Shen, Dajun Zhang · May 10, 2026 · Citations: 0
Rubric-Grounded RL: Structured Judge Rewards for Generalizable Reasoning
Manish Bhattarai, Ismael Boureima, Nishath Rajiv Ranasinghe, Scott Pakin, Dan O'Malley · May 8, 2026 · Citations: 0
RVPO: Risk-Sensitive Alignment via Variance Regularization
Ivan Montero, Tomasz Jurczyk, Bhuwan Dhingra · May 7, 2026 · Citations: 0
RAG over Thinking Traces Can Improve Reasoning Tasks
Negar Arabzadeh, Wenjie Ma, Sewon Min, Matei Zaharia · May 5, 2026 · Citations: 0
Learning to Communicate: Toward End-to-End Optimization of Multi-Agent Language Systems
Ye Yu, Heming Liu, Haibo Jin, Xiaopeng Yuan, Peng Kuang · Apr 23, 2026 · Citations: 0
Process Supervision via Verbal Critique Improves Reasoning in Large Language Models
Hao-Yuan Chen · Apr 23, 2026 · Citations: 0
TRACES: Tagging Reasoning Steps for Adaptive Cost-Efficient Early-Stopping
Yannis Belkhiter, Seshu Tirupathi, Giulio Zizzo, John D. Kelleher · Apr 22, 2026 · Citations: 0
MoE-nD: Per-Layer Mixture-of-Experts Routing for Multi-Axis KV Cache Compression
Libo Sun, Peixiong He, Po-Wei Harn, Xiao Qin · Apr 20, 2026 · Citations: 0
Peer-Predictive Self-Training for Language Model Reasoning
Shi Feng, Hanlin Zhang, Fan Nie, Sham Kakade, Yiling Chen · Apr 14, 2026 · Citations: 0
Sensitivity-Positional Co-Localization in GQA Transformers
Manoj Chandrashekar Rao · Apr 9, 2026 · Citations: 0
Squeeze Evolve: Unified Multi-Model Orchestration for Verifier-Free Evolution
Monishwaran Maheswaran, Leon Lakhani, Zhongzhu Zhou, Shijia Yang, Junxiong Wang · Apr 9, 2026 · Citations: 0
SortedRL: Accelerating RL Training for LLMs through Online Length-Aware Scheduling
Yiqi Zhang, Huiqiang Jiang, Xufang Luo, Zhihe Yang, Chengruidong Zhang · Mar 24, 2026 · Citations: 0
Off-Policy Value-Based Reinforcement Learning for Large Language Models
Peng-Yuan Wang, Ziniu Li, Tian Xu, Bohan Yang, Tian-Shuo Liu · Mar 24, 2026 · Citations: 0
Lie to Me: How Faithful Is Chain-of-Thought Reasoning in Reasoning Models?
Richard J. Young · Mar 23, 2026 · Citations: 0
TERMINATOR: Learning Optimal Exit Points for Early Stopping in Chain-of-Thought Reasoning
Alliot Nagle, Jakhongir Saydaliev, Dhia Garbaya, Michael Gastpar, Ashok Vardhan Makkuva · Mar 13, 2026 · Citations: 0
Tool Verification for Test-Time Reinforcement Learning
Ruotong Liao, Nikolai Röhrich, Xiaohan Wang, Yuhui Zhang, Yasaman Samadzadeh · Mar 2, 2026 · Citations: 0
CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning
Xinyu Zhu, Yihao Feng, Yanchao Sun, Xianzhi Du, Pingzhi Li · Mar 1, 2026 · Citations: 0
Draft-Thinking: Learning Efficient Reasoning in Long Chain-of-Thought LLMs
Jie Cao, Tianwei Lin, Zhenxuan Fan, Bo Yuan, Ziyuan Zhao · Feb 28, 2026 · Citations: 0
LLM Compression by Block Removal with Constrained Binary Optimization
David Jansen, Roman Rausch, Ali Hashemi, David Montero, Román Orús · Jan 29, 2026 · Citations: 0
TRIM: Hybrid Inference via Targeted Stepwise Routing in Multi-Step Reasoning Tasks
Vansh Kapoor, Aman Gupta, Hao Chen, Anurag Beniwal, Jing Huang · Jan 15, 2026 · Citations: 0
PILOT: Planning via Internalized Latent Optimization Trajectories for Large Language Models
Haoyu Zheng, Yun Zhu, Yuqian Yuan, Bo Yuan, Wenqiao Zhang · Jan 7, 2026 · Citations: 0
Self-Consistency Is Losing Its Edge: Diminishing Returns and Rising Costs in Modern LLMs
Chiyan Loo · Nov 2, 2025 · Citations: 0
Merlin's Whisper: Enabling Efficient Reasoning in Large Language Models via Black-box Persuasive Prompting
Heming Xia, Cunxiao Du, Rui Li, Chak Tou Leong, Yongqi Li · Oct 12, 2025 · Citations: 0
SPG: Sandwiched Policy Gradient for Masked Diffusion Language Models
Chenyu Wang, Paria Rashidinejad, DiJia Su, Song Jiang, Sid Wang · Oct 10, 2025 · Citations: 0
Top-H Decoding: Adapting the Creativity and Coherence with Bounded Entropy in Text Generation
Erfan Baghaei Potraghloo, Seyedarmin Azizi, Souvik Kundu, Massoud Pedram · Sep 2, 2025 · Citations: 0
PiCSAR: Probabilistic Confidence Selection And Ranking for Reasoning Chains
Joshua Ong Jun Leang, Zheng Zhao, Aryo Pradipta Gema, Sohee Yang, Wai-Chung Kwan · Aug 29, 2025 · Citations: 0
Strategic Scaling of Test-Time Compute: A Bandit Learning Approach
Bowen Zuo, Yinglun Zhu · Jun 15, 2025 · Citations: 0

Related Benchmark Hubs

DROP Or MATH-500 Or GPQA Benchmark Papers Reasoning & Math Suite Benchmark Papers Reasoning & Math Suite Benchmark Papers In CS.CL MATH-500 Benchmark Papers AIME Or MATH-500 Benchmark Papers GPQA Benchmark Papers GPQA Or BFCL Benchmark Papers (23) HumanEval+ Or BFCL Benchmark Papers (21) MATH-500 Or BFCL Benchmark Papers (25) GPQA Or HumanEval+ Benchmark Papers (22) MATH-500 Or HumanEval+ Benchmark Papers (24) MATH-500 Benchmark Papers (15) MMLU Or AIME Or AlpacaEval Benchmark Papers (60) MMLU Or AIME Or BFCL Benchmark Papers (59) GSM8K Or MMLU Or AIME Benchmark Papers (72) MMLU Or AIME Or HotpotQA Benchmark Papers (62) MMLU Or AIME Or HumanEval+ Benchmark Papers (57)