HFEPX Benchmark Hub

GPQA Or HumanEval+ Benchmark Papers

Updated from current HFEPX corpus (Jun 30, 2026). 41 papers are grouped in this benchmark page.

Read Full Context

Updated from current HFEPX corpus (Jun 30, 2026). 41 papers are grouped in this benchmark page. Common evaluation modes: Automatic Metrics, Llm As Judge. Common annotation unit: Trajectory. Frequently cited benchmark: GPQA. Common metric signal: accuracy. Use this page to compare protocol setup, judge behavior, and labeling design decisions before running new eval experiments. Newest paper in this set is from Apr 1, 2026.

Papers: 41 Last published: Apr 1, 2026 Global RSS

Researcher Quick Triage

Use this page for benchmark-matched method comparisons and eval protocol selection. Quality band: Medium .

High-Signal Coverage

100.0%

41 / 41 sampled papers are not low-signal flagged.

Replication-Ready Set

Papers with explicit benchmark + metric + eval mode fields.

Quality Controls

0.0%

0 papers report calibration/adjudication/IAA controls.

9 papers explicitly name benchmark datasets in the sampled set.
7 papers report at least one metric term in metadata extraction.
Start with the ranked shortlist below before reading all papers.

Primary action: Start with the top 2 benchmark-matched papers, then compare evaluation modes in the protocol matrix.

Why This Matters (Expanded)

Why This Matters For Eval Research

4.9% of papers report explicit human-feedback signals, led by demonstration data.
automatic metrics appears in 17.1% of papers in this hub.
GPQA is a recurring benchmark anchor for cross-paper comparisons in this page.

Protocol Notes (Expanded)

Protocol Takeaways

Quality-control reporting is sparse in this slice; prioritize papers with explicit calibration or adjudication steps.
Rater context is mostly unspecified rater pools, and annotation is commonly trajectory-level annotation; use this to scope replication staffing.
Pair this hub with a human_eval-heavy hub to validate judge-model calibration.

Benchmark Interpretation

GPQA appears in 58.5% of hub papers (24/41); use this cohort for benchmark-matched comparisons.
HumanEval+ appears in 46.3% of hub papers (19/41); use this cohort for benchmark-matched comparisons.

Metric Interpretation

accuracy is reported in 34.1% of hub papers (14/41); compare with a secondary metric before ranking methods.
cost is reported in 19.5% of hub papers (8/41); compare with a secondary metric before ranking methods.

Start Here (Benchmark-Matched First 6)

Ranked by protocol completeness so you can quickly find papers suitable for comparison studies.

S0 Tuning: Zero-Overhead Adaptation of Hybrid Recurrent-Attention Models
Apr 1, 2026 · Citations: 0 · Score: 6.0

Eval: Automatic Metrics · Metrics: Pass@1
Top-b: Entropic Regulation of Relative Probability Bands in Autoregressive Language Processes
Mar 15, 2026 · Citations: 0 · Score: 6.0

Eval: Automatic Metrics · Metrics: Accuracy
D-COT: Disciplined Chain-of-Thought Learning for Efficient Reasoning in Small Language Models
Feb 25, 2026 · Citations: 0 · Score: 6.0

Eval: Automatic Metrics · Metrics: Accuracy
DeepPrune: Parallel Scaling without Inter-trace Redundancy
Oct 9, 2025 · Citations: 0 · Score: 5.5

Eval: Llm As Judge, Automatic Metrics · Metrics: Accuracy
Cost-Effective Communication: An Auction-based Method for Language Agent Interaction
Nov 17, 2025 · Citations: 0 · Score: 5.5

Eval: Automatic Metrics · Metrics: Pass@1
SIGMA: Search-Augmented On-Demand Knowledge Integration for Agentic Mathematical Reasoning
Oct 31, 2025 · Citations: 0 · Score: 5.5

Eval: Automatic Metrics · Metrics: Accuracy

Protocol Matrix (Top 10)

Compare protocol ingredients quickly before deep-reading full papers.

Paper	Eval Modes	Human Feedback	Metrics	Quality Controls
S0 Tuning: Zero-Overhead Adaptation of Hybrid Recurrent-Attention Models Apr 1, 2026	Automatic Metrics	Not reported	Pass@1, Cost	Not reported
Top-b: Entropic Regulation of Relative Probability Bands in Autoregressive Language Processes Mar 15, 2026	Automatic Metrics	Not reported	Accuracy	Not reported
D-COT: Disciplined Chain-of-Thought Learning for Efficient Reasoning in Small Language Models Feb 25, 2026	Automatic Metrics	Not reported	Accuracy	Not reported
DeepPrune: Parallel Scaling without Inter-trace Redundancy Oct 9, 2025	Llm As Judge, Automatic Metrics	Not reported	Accuracy, Auroc	Not reported
Cost-Effective Communication: An Auction-based Method for Language Agent Interaction Nov 17, 2025	Automatic Metrics	Not reported	Pass@1, Cost	Not reported
SIGMA: Search-Augmented On-Demand Knowledge Integration for Agentic Mathematical Reasoning Oct 31, 2025	Automatic Metrics	Not reported	Accuracy	Not reported
Schema for In-Context Learning Oct 14, 2025	Not reported	Demonstrations	Not reported	Not reported
ReCode: Reinforcing Code Generation with Reasoning-Process Rewards Aug 7, 2025	Not reported	Pairwise Preference	Not reported	Not reported
Accelerated Test-Time Scaling with Model-Free Speculative Sampling Jun 5, 2025	Automatic Metrics	Not reported	Accuracy, Latency	Not reported
EntMTP: Accelerating LLM Inference with Entropy Guided Multi Token Prediction Jun 25, 2026	Not reported	Not reported	Not reported	Not reported

Researcher Workflow (Detailed)

Checklist

Gap: Papers with explicit human feedback

Coverage is a replication risk (4.9% vs 45% target).
Gap: Papers reporting quality controls

Coverage is a replication risk (0% vs 30% target).
Strong: Papers naming benchmarks/datasets

Coverage is strong (100% vs 35% target).
Strong: Papers naming evaluation metrics

Coverage is strong (65.9% vs 35% target).
Gap: Papers with known rater population

Coverage is a replication risk (0% vs 35% target).
Gap: Papers with known annotation unit

Coverage is a replication risk (9.8% vs 35% target).

Strengths

Most papers provide measurable evaluation context (100% benchmarks, 65.9% metrics).

Known Gaps

Only 0% of papers report quality controls; prioritize calibration/adjudication evidence.
Rater population is under-specified (0% coverage).
Annotation unit is under-specified (9.8% coverage).

Suggested Next Analyses

Pair this hub with a human_eval-heavy hub to validate judge-model calibration.
Stratify by benchmark (GPQA vs HumanEval+) before comparing methods.
Track metric sensitivity by reporting both accuracy and cost.

Recommended Queries

LLM-as-Judge Protocols Benchmark Slice: GPQA Metric Slice: accuracy Recent High-Signal Papers

Known Limitations

Only 0% of papers report quality controls; prioritize calibration/adjudication evidence.
Rater population is under-specified (0% coverage).
Narrative synthesis is grounded in metadata and abstracts only; full-paper implementation details are not parsed.

Research Utility Snapshot (Detailed)

Evaluation Modes

Automatic Metrics (7)
Llm As Judge (1)

Human Feedback Mix

Demonstrations (1)
Pairwise Preference (1)

Top Benchmarks

GPQA (24)
HumanEval+ (19)
GSM8K (12)
AIME (10)

Top Metrics

Accuracy (14)
Cost (8)
Pass@1 (3)
Coherence (2)

Top Papers On This Benchmark

S0 Tuning: Zero-Overhead Adaptation of Hybrid Recurrent-Attention Models
Jack Young · Apr 1, 2026 · Citations: 0

Automatic Metrics

Using roughly 48 execution-verified HumanEval training solutions, tuning a single initial state matrix per recurrent layer, with zero inference overhead, outperforms LoRA by +10.8 pp (p < 0.001) on HumanEval.
Top-b: Entropic Regulation of Relative Probability Bands in Autoregressive Language Processes
Deepon Halder, Raj Dabre · Mar 15, 2026 · Citations: 0

Automatic Metrics

Empirical validation on GPQA and GSM8K benchmarks indicates that Top-b significantly reduces generation entropy and inter-decoding variance while maintaining competitive reasoning accuracy, effectively approximating a self-regulating…
D-COT: Disciplined Chain-of-Thought Learning for Efficient Reasoning in Small Language Models
Shunsuke Ubukata · Feb 25, 2026 · Citations: 0

Automatic Metrics

In this study, we propose Disciplined Chain-of-Thought (D-CoT), a novel framework that enforces a structured reasoning process using control tags -- such as <TEMP_LOW> for fact-checking and <TEMP_HIGH> for multi-perspective exploration --…
Accelerated Test-Time Scaling with Model-Free Speculative Sampling
Woomin Song, Saket Dingliwal, Sai Muralidhar Jayanthi, Bhavana Ganesh, Jinwoo Shin · Jun 5, 2025 · Citations: 0

Automatic Metrics

Extensive evaluations across multiple models and reasoning tasks (AIME-2024, GPQA-Diamond, and LiveCodeBench) demonstrate that STAND reduces inference latency by 60-65% compared to standard autoregressive decoding while maintaining…
DeepPrune: Parallel Scaling without Inter-trace Redundancy
Shangqing Tu, Yaxuan Li, Yushi Bai, Lei Hou, Juanzi Li · Oct 9, 2025 · Citations: 0

Llm As JudgeAutomatic Metrics

Our method features a specialized judge model trained with out-of-distribution data (AIME 2022, AIME 2023, and MATH 500) using oversampling techniques to accurately predict answer equivalence from partial reasoning traces, achieving 0.7072…
Cost-Effective Communication: An Auction-based Method for Language Agent Interaction
Yijia Fan, Jusheng Zhang, Kaitong Cai, Jing Yang, Chengpei Tang · Nov 17, 2025 · Citations: 0

Automatic Metrics

To address this, we introduce the Dynamic Auction-based Language Agent (DALA), a novel framework that treats communication bandwidth as a scarce and tradable resource.
SIGMA: Search-Augmented On-Demand Knowledge Integration for Agentic Mathematical Reasoning
Ali Asgarov, Umid Suleymanov, Aadyant Khatri · Oct 31, 2025 · Citations: 0

Automatic Metrics

We introduce SIGMA (Search-Augmented On-Demand Knowledge Integration for AGentic Mathematical reAsoning), a unified framework that orchestrates specialized agents to independently reason, perform targeted searches, and synthesize findings…
Schema for In-Context Learning
Pan Chen, Shaohong Chen, Mark Wang, Shi Xuan Leong, Priscilla Fung · Oct 14, 2025 · Citations: 0

Demonstrations

Inspired by cognitive science, specifically schema theory, which holds that humans interpret new information by activating pre-existing mental frameworks (schemas) to structure understanding, we introduce Schema-Activated In-Context…
ReCode: Reinforcing Code Generation with Reasoning-Process Rewards
Lishui Fan, Yu Zhang, Mouxiang Chen, Zhongxin Liu · Aug 7, 2025 · Citations: 0

Pairwise Preference

Additionally, to assess the reward model's discriminative capability in assessing reasoning-process quality, we introduce LiveCodeBench-RewardBench (LCB-RB), a new benchmark comprising preference pairs of superior and inferior reasoning…
EntMTP: Accelerating LLM Inference with Entropy Guided Multi Token Prediction
Carrie Chen · Jun 25, 2026 · Citations: 0
BrahmicTokenizer-131K: An Indic-Capable Drop-In Replacement for o200k_base
Rohan Shravan · May 28, 2026 · Citations: 0
ACC: Compiling Agent Trajectories for Long-Context Training
Qisheng Su, Zhen Fang, Shiting Huang, Yu Zeng, Yiming Zhao · May 21, 2026 · Citations: 0
LamPO: A Lambda Style Policy Optimization for Reasoning Language Models
Zhe Yuan, Yipeng Zhou, Jinghan Li, Xinyuan Chen, Bowen Deng · May 20, 2026 · Citations: 0
FocuSFT: Bilevel Optimization for Dilution-Aware Long-Context Fine-Tuning
Zehua Pei, Hui-Ling Zhen, Xianzhi Yu, Sinno Jialin Pan, Mingxuan Yuan · May 11, 2026 · Citations: 0
The Metacognitive Probe: Five Behavioural Calibration Diagnostics for LLMs
Rafael C. T. Oliveira · May 11, 2026 · Citations: 0
Parameter-Efficient Neuroevolution for Diverse LLM Generation: Quality-Diversity Optimization via Prompt Embedding Evolution
Dongxin Guo, Jikun Wu, Siu Ming Yiu · May 10, 2026 · Citations: 0
Edit-Based Refinement for Parallel Masked Diffusion Language Models
Houxing Ren, Mingjie Zhan, Zimu Lu, Ke Wang, Yunqiao Yang · May 10, 2026 · Citations: 0
Rubric-Grounded RL: Structured Judge Rewards for Generalizable Reasoning
Manish Bhattarai, Ismael Boureima, Nishath Rajiv Ranasinghe, Scott Pakin, Dan O'Malley · May 8, 2026 · Citations: 0
Unsolvability Ceiling in Multi-LLM Routing: An Empirical Study of Evaluation Artifacts
Saloni Garg, Amit Sagtani · May 8, 2026 · Citations: 0
RVPO: Risk-Sensitive Alignment via Variance Regularization
Ivan Montero, Tomasz Jurczyk, Bhuwan Dhingra · May 7, 2026 · Citations: 0
RAG over Thinking Traces Can Improve Reasoning Tasks
Negar Arabzadeh, Wenjie Ma, Sewon Min, Matei Zaharia · May 5, 2026 · Citations: 0
Turning the TIDE: Cross-Architecture Distillation for Diffusion Large Language Models
Gongbo Zhang, Wen Wang, Ye Tian, Li Yuan · Apr 29, 2026 · Citations: 0
Learning to Communicate: Toward End-to-End Optimization of Multi-Agent Language Systems
Ye Yu, Heming Liu, Haibo Jin, Xiaopeng Yuan, Peng Kuang · Apr 23, 2026 · Citations: 0
Process Supervision via Verbal Critique Improves Reasoning in Large Language Models
Hao-Yuan Chen · Apr 23, 2026 · Citations: 0
TRACES: Tagging Reasoning Steps for Adaptive Cost-Efficient Early-Stopping
Yannis Belkhiter, Seshu Tirupathi, Giulio Zizzo, John D. Kelleher · Apr 22, 2026 · Citations: 0
RACER: Retrieval-Augmented Contextual Rapid Speculative Decoding
Zihong Zhang, Zuchao Li, Lefei Zhang, Ping Wang, Hai Zhao · Apr 16, 2026 · Citations: 0
StoryCoder: Narrative Reformulation for Structured Reasoning in LLM Code Generation
Geonhui Jang, Dongyoon Han, YoungJoon Yoo · Apr 16, 2026 · Citations: 0
Sensitivity-Positional Co-Localization in GQA Transformers
Manoj Chandrashekar Rao · Apr 9, 2026 · Citations: 0
Squeeze Evolve: Unified Multi-Model Orchestration for Verifier-Free Evolution
Monishwaran Maheswaran, Leon Lakhani, Zhongzhu Zhou, Shijia Yang, Junxiong Wang · Apr 9, 2026 · Citations: 0
Off-Policy Value-Based Reinforcement Learning for Large Language Models
Peng-Yuan Wang, Ziniu Li, Tian Xu, Bohan Yang, Tian-Shuo Liu · Mar 24, 2026 · Citations: 0
Lie to Me: How Faithful Is Chain-of-Thought Reasoning in Reasoning Models?
Richard J. Young · Mar 23, 2026 · Citations: 0
TERMINATOR: Learning Optimal Exit Points for Early Stopping in Chain-of-Thought Reasoning
Alliot Nagle, Jakhongir Saydaliev, Dhia Garbaya, Michael Gastpar, Ashok Vardhan Makkuva · Mar 13, 2026 · Citations: 0
In-Context Environments Induce Evaluation-Awareness in Language Models
Maheep Chaudhary · Mar 4, 2026 · Citations: 0
CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning
Xinyu Zhu, Yihao Feng, Yanchao Sun, Xianzhi Du, Pingzhi Li · Mar 1, 2026 · Citations: 0
Distribution-Aware Companding Quantization of Large Language Models
Athul Radhakrishnan, Siddhant Mohan, Mahima Sachdeva · Feb 27, 2026 · Citations: 0
LLM Compression by Block Removal with Constrained Binary Optimization
David Jansen, Roman Rausch, Ali Hashemi, David Montero, Román Orús · Jan 29, 2026 · Citations: 0
Bridging Draft Policy Misalignment: Group Tree Optimization for Speculative Decoding
Shijing Hu, Jingyang Li, Zhihui Lu, Pan Zhou · Sep 26, 2025 · Citations: 0
Top-H Decoding: Adapting the Creativity and Coherence with Bounded Entropy in Text Generation
Erfan Baghaei Potraghloo, Seyedarmin Azizi, Souvik Kundu, Massoud Pedram · Sep 2, 2025 · Citations: 0
SafeSieve: From Heuristics to Experience in Progressive Pruning for LLM-based Multi-Agent Communication
Ruijia Zhang, Xinyan Zhao, Ruixiang Wang, Sigen Chen, Guibin Zhang · Aug 15, 2025 · Citations: 0
Is Human-Like Text Liked by Humans? Multilingual Human Detection and Preference Against AI
Yuxia Wang, Rui Xing, Jonibek Mansurov, Giovanni Puccetti, Zhuohan Xie · Feb 17, 2025 · Citations: 0
LoRA-FA: Efficient and Effective Low Rank Representation Fine-tuning
Longteng Zhang, Lin Zhang, Shaohuai Shi, Xiaowen Chu, Bo Li · Aug 7, 2023 · Citations: 0

Related Benchmark Hubs

DROP Or GPQA Or HumanEval+ Benchmark Papers GPQA Benchmark Papers GPQA In CS.CL Papers GPQA Benchmark Papers (Last 365 Days) GPQA Benchmark Papers (Last 300 Days) GSM8K Or GPQA Benchmark Papers GPQA Or BFCL Benchmark Papers (23) HumanEval+ Or BFCL Benchmark Papers (21) MATH-500 Or BFCL Benchmark Papers (25) MATH-500 Or GPQA Benchmark Papers (25) MATH-500 Or HumanEval+ Benchmark Papers (24) MATH-500 Benchmark Papers (15) MMLU Or AIME Or AlpacaEval Benchmark Papers (60) MMLU Or AIME Or BFCL Benchmark Papers (59) GSM8K Or MMLU Or AIME Benchmark Papers (72) MMLU Or AIME Or HotpotQA Benchmark Papers (62) MMLU Or AIME Or HumanEval+ Benchmark Papers (57)