HFEPX Benchmark Hub

GPQA Or BFCL Benchmark Papers

Updated from current HFEPX corpus (Jun 30, 2026). 37 papers are grouped in this benchmark page.

Read Full Context

Updated from current HFEPX corpus (Jun 30, 2026). 37 papers are grouped in this benchmark page. Common evaluation modes: Automatic Metrics, Llm As Judge. Common annotation unit: Trajectory. Frequently cited benchmark: GPQA. Common metric signal: accuracy. Use this page to compare protocol setup, judge behavior, and labeling design decisions before running new eval experiments. Newest paper in this set is from Apr 2, 2026.

Papers: 37 Last published: Apr 2, 2026 Global RSS

Researcher Quick Triage

Use this page for benchmark-matched method comparisons and eval protocol selection. Quality band: Medium .

High-Signal Coverage

100.0%

37 / 37 sampled papers are not low-signal flagged.

Replication-Ready Set

Papers with explicit benchmark + metric + eval mode fields.

Quality Controls

0.0%

0 papers report calibration/adjudication/IAA controls.

10 papers explicitly name benchmark datasets in the sampled set.
9 papers report at least one metric term in metadata extraction.
Start with the ranked shortlist below before reading all papers.

Primary action: Start with the top 2 benchmark-matched papers, then compare evaluation modes in the protocol matrix.

Why This Matters (Expanded)

Why This Matters For Eval Research

2.7% of papers report explicit human-feedback signals, led by demonstration data.
automatic metrics appears in 24.3% of papers in this hub.
GPQA is a recurring benchmark anchor for cross-paper comparisons in this page.

Protocol Notes (Expanded)

Protocol Takeaways

Quality-control reporting is sparse in this slice; prioritize papers with explicit calibration or adjudication steps.
Rater context is mostly unspecified rater pools, and annotation is commonly trajectory-level annotation; use this to scope replication staffing.
Pair this hub with a human_eval-heavy hub to validate judge-model calibration.

Benchmark Interpretation

GPQA appears in 64.9% of hub papers (24/37); use this cohort for benchmark-matched comparisons.
BFCL appears in 35.1% of hub papers (13/37); use this cohort for benchmark-matched comparisons.

Metric Interpretation

accuracy is reported in 37.8% of hub papers (14/37); compare with a secondary metric before ranking methods.
cost is reported in 16.2% of hub papers (6/37); compare with a secondary metric before ranking methods.

Start Here (Benchmark-Matched First 6)

Ranked by protocol completeness so you can quickly find papers suitable for comparison studies.

Brief Is Better: Non-Monotonic Chain-of-Thought Budget Effects in Function-Calling Language Agents
Apr 2, 2026 · Citations: 0 · Score: 6.0

Eval: Automatic Metrics · Metrics: Accuracy
Top-b: Entropic Regulation of Relative Probability Bands in Autoregressive Language Processes
Mar 15, 2026 · Citations: 0 · Score: 6.0

Eval: Automatic Metrics · Metrics: Accuracy
D-COT: Disciplined Chain-of-Thought Learning for Efficient Reasoning in Small Language Models
Feb 25, 2026 · Citations: 0 · Score: 6.0

Eval: Automatic Metrics · Metrics: Accuracy
SkillX: Automatically Constructing Skill Knowledge Bases for Agents
Apr 6, 2026 · Citations: 0 · Score: 6.0

Eval: Automatic Metrics · Metrics: Task success
DeepPrune: Parallel Scaling without Inter-trace Redundancy
Oct 9, 2025 · Citations: 0 · Score: 5.5

Eval: Llm As Judge, Automatic Metrics · Metrics: Accuracy
The Bitter Lesson of Diffusion Language Models for Agentic Workflows: A Comprehensive Reality Check
Jan 19, 2026 · Citations: 0 · Score: 5.5

Eval: Automatic Metrics · Metrics: Precision

Protocol Matrix (Top 10)

Compare protocol ingredients quickly before deep-reading full papers.

Paper	Eval Modes	Human Feedback	Metrics	Quality Controls
Brief Is Better: Non-Monotonic Chain-of-Thought Budget Effects in Function-Calling Language Agents Apr 2, 2026	Automatic Metrics	Not reported	Accuracy	Not reported
Top-b: Entropic Regulation of Relative Probability Bands in Autoregressive Language Processes Mar 15, 2026	Automatic Metrics	Not reported	Accuracy	Not reported
D-COT: Disciplined Chain-of-Thought Learning for Efficient Reasoning in Small Language Models Feb 25, 2026	Automatic Metrics	Not reported	Accuracy	Not reported
SkillX: Automatically Constructing Skill Knowledge Bases for Agents Apr 6, 2026	Automatic Metrics	Not reported	Task success	Not reported
DeepPrune: Parallel Scaling without Inter-trace Redundancy Oct 9, 2025	Llm As Judge, Automatic Metrics	Not reported	Accuracy, Auroc	Not reported
The Bitter Lesson of Diffusion Language Models for Agentic Workflows: A Comprehensive Reality Check Jan 19, 2026	Automatic Metrics	Not reported	Precision, Latency	Not reported
SIGMA: Search-Augmented On-Demand Knowledge Integration for Agentic Mathematical Reasoning Oct 31, 2025	Automatic Metrics	Not reported	Accuracy	Not reported
Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions Sep 23, 2025	Automatic Metrics	Not reported	Accuracy	Not reported
Schema for In-Context Learning Oct 14, 2025	Not reported	Demonstrations	Not reported	Not reported
Accelerated Test-Time Scaling with Model-Free Speculative Sampling Jun 5, 2025	Automatic Metrics	Not reported	Accuracy, Latency	Not reported

Researcher Workflow (Detailed)

Checklist

Gap: Papers with explicit human feedback

Coverage is a replication risk (2.7% vs 45% target).
Gap: Papers reporting quality controls

Coverage is a replication risk (0% vs 30% target).
Strong: Papers naming benchmarks/datasets

Coverage is strong (100% vs 35% target).
Strong: Papers naming evaluation metrics

Coverage is strong (67.6% vs 35% target).
Gap: Papers with known rater population

Coverage is a replication risk (0% vs 35% target).
Gap: Papers with known annotation unit

Coverage is a replication risk (10.8% vs 35% target).

Strengths

Most papers provide measurable evaluation context (100% benchmarks, 67.6% metrics).

Known Gaps

Only 0% of papers report quality controls; prioritize calibration/adjudication evidence.
Rater population is under-specified (0% coverage).
Annotation unit is under-specified (10.8% coverage).

Suggested Next Analyses

Pair this hub with a human_eval-heavy hub to validate judge-model calibration.
Stratify by benchmark (GPQA vs BFCL) before comparing methods.
Track metric sensitivity by reporting both accuracy and cost.

Recommended Queries

LLM-as-Judge Protocols Benchmark Slice: GPQA Metric Slice: accuracy Recent High-Signal Papers

Known Limitations

Only 0% of papers report quality controls; prioritize calibration/adjudication evidence.
Rater population is under-specified (0% coverage).
Narrative synthesis is grounded in metadata and abstracts only; full-paper implementation details are not parsed.

Research Utility Snapshot (Detailed)

Evaluation Modes

Automatic Metrics (9)
Llm As Judge (1)

Human Feedback Mix

Demonstrations (1)

Top Benchmarks

GPQA (24)
BFCL (13)
AIME (11)
MMLU (7)

Top Metrics

Accuracy (14)
Cost (6)
Latency (2)
Spearman (2)

Top Papers On This Benchmark

Brief Is Better: Non-Monotonic Chain-of-Thought Budget Effects in Function-Calling Language Agents
Xuan Qi · Apr 2, 2026 · Citations: 0

Automatic Metrics

Chain-of-thought (CoT) reasoning is widely assumed to improve agent performance, but the relationship between reasoning length and accuracy in structured tool-use settings remains poorly understood.
Top-b: Entropic Regulation of Relative Probability Bands in Autoregressive Language Processes
Deepon Halder, Raj Dabre · Mar 15, 2026 · Citations: 0

Automatic Metrics

Empirical validation on GPQA and GSM8K benchmarks indicates that Top-b significantly reduces generation entropy and inter-decoding variance while maintaining competitive reasoning accuracy, effectively approximating a self-regulating…
D-COT: Disciplined Chain-of-Thought Learning for Efficient Reasoning in Small Language Models
Shunsuke Ubukata · Feb 25, 2026 · Citations: 0

Automatic Metrics

In this study, we propose Disciplined Chain-of-Thought (D-CoT), a novel framework that enforces a structured reasoning process using control tags -- such as <TEMP_LOW> for fact-checking and <TEMP_HIGH> for multi-perspective exploration --…
Accelerated Test-Time Scaling with Model-Free Speculative Sampling
Woomin Song, Saket Dingliwal, Sai Muralidhar Jayanthi, Bhavana Ganesh, Jinwoo Shin · Jun 5, 2025 · Citations: 0

Automatic Metrics

Extensive evaluations across multiple models and reasoning tasks (AIME-2024, GPQA-Diamond, and LiveCodeBench) demonstrate that STAND reduces inference latency by 60-65% compared to standard autoregressive decoding while maintaining…
DeepPrune: Parallel Scaling without Inter-trace Redundancy
Shangqing Tu, Yaxuan Li, Yushi Bai, Lei Hou, Juanzi Li · Oct 9, 2025 · Citations: 0

Llm As JudgeAutomatic Metrics

Our method features a specialized judge model trained with out-of-distribution data (AIME 2022, AIME 2023, and MATH 500) using oversampling techniques to accurately predict answer equivalence from partial reasoning traces, achieving 0.7072…
SkillX: Automatically Constructing Skill Knowledge Bases for Agents
Chenxi Wang, Zhuoyun Yu, Xin Xie, Wuguannan Yao, Runnan Fang · Apr 6, 2026 · Citations: 0

Automatic Metrics

Learning from experience is critical for building capable large language model (LLM) agents, yet prevailing self-evolving paradigms remain inefficient: agents learn in isolation, repeatedly rediscover similar behaviors from limited…
The Bitter Lesson of Diffusion Language Models for Agentic Workflows: A Comprehensive Reality Check
Qingyu Lu, Liang Ding, Kanjian Zhang, Jinxia Zhang, Dacheng Tao · Jan 19, 2026 · Citations: 0

Automatic Metrics

In this work, we present a comprehensive evaluation of dLLMs (e.g., LLaDA, Dream) across two distinct agentic paradigms: Embodied Agents (requiring long-horizon planning) and Tool-Calling Agents (requiring precise formatting).
SIGMA: Search-Augmented On-Demand Knowledge Integration for Agentic Mathematical Reasoning
Ali Asgarov, Umid Suleymanov, Aadyant Khatri · Oct 31, 2025 · Citations: 0

Automatic Metrics

We introduce SIGMA (Search-Augmented On-Demand Knowledge Integration for AGentic Mathematical reAsoning), a unified framework that orchestrates specialized agents to independently reason, perform targeted searches, and synthesize findings…
Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions
Junhao Su, Yuanliang Wan, Junwei Yang, Hengyu Shi, Tianyang Han · Sep 23, 2025 · Citations: 0

Automatic Metrics

The agent produces a short yet precise reflection: it diagnoses the failure using evidence from the previous step and then proposes a correct, executable follow-up call.
Schema for In-Context Learning
Pan Chen, Shaohong Chen, Mark Wang, Shi Xuan Leong, Priscilla Fung · Oct 14, 2025 · Citations: 0

Demonstrations

Inspired by cognitive science, specifically schema theory, which holds that humans interpret new information by activating pre-existing mental frameworks (schemas) to structure understanding, we introduce Schema-Activated In-Context…
Notation Matters: A Benchmark Study of Token-Optimized Formats in Agentic AI Systems
Lorenz Kutschka, Bernhard Geiger · May 28, 2026 · Citations: 0
ACC: Compiling Agent Trajectories for Long-Context Training
Qisheng Su, Zhen Fang, Shiting Huang, Yu Zeng, Yiming Zhao · May 21, 2026 · Citations: 0
LamPO: A Lambda Style Policy Optimization for Reasoning Language Models
Zhe Yuan, Yipeng Zhou, Jinghan Li, Xinyuan Chen, Bowen Deng · May 20, 2026 · Citations: 0
HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agents
Woongyeng Yeo, Yumin Choi, Taekyung Ki, Sung Ju Hwang · May 18, 2026 · Citations: 0
TIER: Trajectory-Invariant Execution Rewards for Multi-Step Tool Composition
Anay Kulkarni, ChiaEn Lu, Dheeraj Mekala, Jayanth Srinivasa, Gaowen Liu · May 16, 2026 · Citations: 0
FocuSFT: Bilevel Optimization for Dilution-Aware Long-Context Fine-Tuning
Zehua Pei, Hui-Ling Zhen, Xianzhi Yu, Sinno Jialin Pan, Mingxuan Yuan · May 11, 2026 · Citations: 0
The Metacognitive Probe: Five Behavioural Calibration Diagnostics for LLMs
Rafael C. T. Oliveira · May 11, 2026 · Citations: 0
Rubric-Grounded RL: Structured Judge Rewards for Generalizable Reasoning
Manish Bhattarai, Ismael Boureima, Nishath Rajiv Ranasinghe, Scott Pakin, Dan O'Malley · May 8, 2026 · Citations: 0
RVPO: Risk-Sensitive Alignment via Variance Regularization
Ivan Montero, Tomasz Jurczyk, Bhuwan Dhingra · May 7, 2026 · Citations: 0
RAG over Thinking Traces Can Improve Reasoning Tasks
Negar Arabzadeh, Wenjie Ma, Sewon Min, Matei Zaharia · May 5, 2026 · Citations: 0
Learning to Communicate: Toward End-to-End Optimization of Multi-Agent Language Systems
Ye Yu, Heming Liu, Haibo Jin, Xiaopeng Yuan, Peng Kuang · Apr 23, 2026 · Citations: 0
Process Supervision via Verbal Critique Improves Reasoning in Large Language Models
Hao-Yuan Chen · Apr 23, 2026 · Citations: 0
TRACES: Tagging Reasoning Steps for Adaptive Cost-Efficient Early-Stopping
Yannis Belkhiter, Seshu Tirupathi, Giulio Zizzo, John D. Kelleher · Apr 22, 2026 · Citations: 0
Breaking MCP with Function Hijacking Attacks: Novel Threats for Function Calling and Agentic Models
Yannis Belkhiter, Giulio Zizzo, Sergio Maffeis, Seshu Tirupathi, John D. Kelleher · Apr 22, 2026 · Citations: 0
CoEvolve: Training LLM Agents via Agent-Data Mutual Evolution
Shidong Yang, Ziyu Ma, Tongwen Huang, Yiming Hu, Yong Wang · Apr 17, 2026 · Citations: 0
Awakening the Sleeping Agent: Lean-Specific Agentic Data Reactivates General Tool Use in Goedel Prover
Jui-Hui Chung, Hongzhou Lin, Lai Jiang, Shange Tang, Chi Jin · Apr 9, 2026 · Citations: 0
Sensitivity-Positional Co-Localization in GQA Transformers
Manoj Chandrashekar Rao · Apr 9, 2026 · Citations: 0
Squeeze Evolve: Unified Multi-Model Orchestration for Verifier-Free Evolution
Monishwaran Maheswaran, Leon Lakhani, Zhongzhu Zhou, Shijia Yang, Junxiong Wang · Apr 9, 2026 · Citations: 0
Off-Policy Value-Based Reinforcement Learning for Large Language Models
Peng-Yuan Wang, Ziniu Li, Tian Xu, Bohan Yang, Tian-Shuo Liu · Mar 24, 2026 · Citations: 0
Lie to Me: How Faithful Is Chain-of-Thought Reasoning in Reasoning Models?
Richard J. Young · Mar 23, 2026 · Citations: 0
TERMINATOR: Learning Optimal Exit Points for Early Stopping in Chain-of-Thought Reasoning
Alliot Nagle, Jakhongir Saydaliev, Dhia Garbaya, Michael Gastpar, Ashok Vardhan Makkuva · Mar 13, 2026 · Citations: 0
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
Ben Rank, Hardik Bhatnagar, Ameya Prabhu, Shira Eisenberg, Karina Nguyen · Mar 9, 2026 · Citations: 0
CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning
Xinyu Zhu, Yihao Feng, Yanchao Sun, Xianzhi Du, Pingzhi Li · Mar 1, 2026 · Citations: 0
LLM Compression by Block Removal with Constrained Binary Optimization
David Jansen, Roman Rausch, Ali Hashemi, David Montero, Román Orús · Jan 29, 2026 · Citations: 0
Beyond Max Tokens: Stealthy Resource Amplification via Tool Calling Chains in LLM Agents
Kaiyu Zhou, Yongsen Zheng, Yicheng He, Meng Xue, Xueluan Gong · Jan 16, 2026 · Citations: 0
Remember Me, Refine Me: A Dynamic Procedural Memory Framework for Experience-Driven Agent Evolution
Zouying Cao, Jiaji Deng, Li Yu, Weikang Zhou, Zhaoyang Liu · Dec 11, 2025 · Citations: 0
Top-H Decoding: Adapting the Creativity and Coherence with Bounded Entropy in Text Generation
Erfan Baghaei Potraghloo, Seyedarmin Azizi, Souvik Kundu, Massoud Pedram · Sep 2, 2025 · Citations: 0

Related Benchmark Hubs

DROP Or GPQA Or BFCL Benchmark Papers GPQA Benchmark Papers GPQA In CS.CL Papers GPQA Benchmark Papers (Last 365 Days) GPQA Benchmark Papers (Last 300 Days) GPQA Or ToolBench Benchmark Papers HumanEval+ Or BFCL Benchmark Papers (21) MATH-500 Or BFCL Benchmark Papers (25) GPQA Or HumanEval+ Benchmark Papers (22) MATH-500 Or GPQA Benchmark Papers (25) MATH-500 Or HumanEval+ Benchmark Papers (24) MATH-500 Benchmark Papers (15) MMLU Or AIME Or AlpacaEval Benchmark Papers (60) MMLU Or AIME Or BFCL Benchmark Papers (59) GSM8K Or MMLU Or AIME Benchmark Papers (72) MMLU Or AIME Or HotpotQA Benchmark Papers (62) MMLU Or AIME Or HumanEval+ Benchmark Papers (57)