HFEPX Benchmark Hub
BFCL or MATH-500 Benchmark Papers
Updated from current HFEPX corpus (2026-09-13). This page tracks 45 papers reporting BFCL or MATH-500 benchmark evidence, with protocol and metric context for comparison.
HFEPX Benchmark Hub
Updated from current HFEPX corpus (2026-09-13). This page tracks 45 papers reporting BFCL or MATH-500 benchmark evidence, with protocol and metric context for comparison.
Use this page for benchmark-matched method comparisons and eval protocol selection. Quality band: High .
High-Signal Coverage
100.0%
45 / 45 sampled papers are not low-signal flagged.
Replication-Ready Set
28
Papers with explicit benchmark + metric + eval mode fields.
Quality Controls
4.4%
2 papers report calibration/adjudication/IAA controls.
Primary action: Start with the top 2 benchmark-matched papers, then compare evaluation modes in the protocol matrix.
Ranked by protocol completeness so you can quickly find papers suitable for comparison studies.
Jun 24, 2026 · Citations: 0 · Score: 8.0
Eval: Automatic Metrics · Metrics: Accuracy
May 30, 2026 · Citations: 0 · Score: 7.5
Eval: Simulation Env · Metrics: Precision
Apr 22, 2026 · Citations: 0 · Score: 7.5
Eval: Automatic Metrics · Metrics: Jailbreak success rate
Apr 1, 2026 · Citations: 0 · Score: 7.5
Eval: Automatic Metrics · Metrics: Error rate
Sep 27, 2025 · Citations: 0 · Score: 7.0
Eval: Automatic Metrics · Metrics: Accuracy
Nov 3, 2025 · Citations: 0 · Score: 7.0
Eval: Automatic Metrics · Metrics: Accuracy
Compare protocol ingredients quickly before deep-reading full papers.
Gap: Human feedback
Human feedback is present in 4 of 45 papers.
Gap: Quality controls
Quality controls is present in 2 of 45 papers.
Strong: Benchmarks
Benchmarks is present in 45 of 45 papers.
Strong: Metrics
Metrics is present in 29 of 45 papers.
Gap: Known rater population
Known rater population is present in 2 of 45 papers.
Moderate: Known annotation unit
Known annotation unit is present in 10 of 45 papers.
Evaluation Modes
Human Feedback Mix
Top Benchmarks
Top Metrics
Chishui Chen, Jiaye Lin, Te Sun, Yi Yang, Junxi Wang · May 30, 2026 · Citations: 0
Agent skills are callable procedural modules that provide reusable knowledge and execution policies for complex agentic tasks.
Bo Liu, Simon Yu, Yiding Jiang, Ao Qu, Andrew Zhao · Aug 19, 2026 · Citations: 0
For language agents, existing training environment pools (hand-curated, statically synthesized, or frozen-verifier) keep the goal distribution fixed as the learner scales.
Yannis Belkhiter, Giulio Zizzo, Sergio Maffeis, Seshu Tirupathi, John D. Kelleher · Apr 22, 2026 · Citations: 0
The growth of agentic AI has drawn significant attention to function calling Large Language Models (LLMs), which are designed to extend the capabilities of AI-powered system by invoking external functions.
Ruotong Liao, Nikolai Röhrich, Xiaohan Wang, Yuhui Zhang, Yasaman Samadzadeh · Mar 2, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Alessio Bruno · May 30, 2026 · Citations: 0
The rule-only path answers the 20,000-record lm-eval arithmetic benchmark at 100%, 1 ms per record.
Dylan Bouchard, Mohit Singh Chauhan · Aug 12, 2026 · Citations: 0
For LLM agents, however, the unit of observation is an interactive trajectory, where the model can ask clarifying questions, call tools, update state, and make intermediate decisions whose errors propagate to the final outcome.
Zhenyu Zhao, Sander Land, Daniel M. Bikel, Waseem Alshikh · Apr 29, 2026 · Citations: 0
Across three model families and five mathematical reasoning benchmarks, our approach shortens reasoning traces by 8.1% on average; under a TOST equivalence analysis at a +/- 2pp margin, accuracy is equivalent or inconclusive on 13/15 model…
Ningning Xu, Yuxuan Jiang, Shubhashis Roy Dipta, Hengyuan Zhang · Sep 27, 2025 · Citations: 0
We propose a two-stage framework that first builds code competence from both patterns and then aligns pattern selection with teacher preferences.
Steven Kolawole, Virginia Smith · Jun 25, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Lanxiang Hu, Zhaoxiang Feng, Yulun Wu, Haoran Yuan, Yujie Zhao · Jun 16, 2026 · Citations: 0
Across math, coding, and chat benchmarks on dense and MoE Qwen3 models, JetSpec consistently outperforms bidirectional-head and tree-based SD baselines.
Jaeyong Ko, Pilsung Kang, Yukyung Lee · Jun 24, 2026 · Citations: 0
Across seven models and three mathematical reasoning benchmarks (GSM1K, MATH500, AIME 2025), cliff tokens act as failure triggers; deleting the first cliff token and resampling recovers pass@64 to 1.0, while keeping it limits recovery to…
Jie Deng, Shining Liang, Jun Li, Hongzhi Li, Yutao Xie · Feb 1, 2026 · Citations: 0
This phenomenon arises from multi-question contextual pressure during generation and consistently manifests across models and benchmarks.
Liwen Zheng, Haiyun Jiang · Jun 23, 2026 · Citations: 0
In a six-variant Qwen3 math reasoning benchmark with a uniform 200-step training budget for all trained variants, we use pass@8 as the primary problem-level solve-rate metric.
Ali Asgarov, Umid Suleymanov, Aadyant Khatri · Oct 31, 2025 · Citations: 0
We introduce SIGMA (Search-Augmented On-Demand Knowledge Integration for AGentic Mathematical reAsoning), a unified framework that orchestrates specialized agents to independently reason, perform targeted searches, and synthesize findings…
Felix Zhou, Anay Mehrotra, Quanquan C. Liu · May 28, 2026 · Citations: 0
Across MATH500, HumanEval, GPQA Diamond, and AIME26, our method consistently improves over baselines and RL-trained models.
G M Shahariar, Erfan Shayegani, Ali Nazari, Nael Abu-Ghazaleh · Oct 25, 2025 · Citations: 0
Experiments across four benchmarks (AIME25, MATH-500, GSM8k, and GPQA Diamond) using three state-of-the-art open reasoning models demonstrate that Q-Value steering policy achieves significant performance gains with "surgical" efficiency,…
Andrea Sassella, Andrea Chizzola, Tommaso Bianchi, Luca Alessandrelli, Mark James Carman · May 8, 2026 · Citations: 0
This report benchmarks the performance of ENGINEERING Ingegneria Informatica S.p.A.'s EngGPT2MoE-16B-A3B LLM, a 16B parameter Mixture of Experts (MoE) model with 3B active parameters.
Hang Yan, Fangzhi Xu, Rongman Xu, Yifei Li, Jian Zhang · Jul 20, 2025 · Citations: 0
MUR is comprehensively evaluated against various TTS methods across four challenging benchmarks (MATH-500, AIME24, AIME25, and GPQA-diamond) using different sizes of recent Qwen3 models (1.7B, 4B, and 8B).
Zekun Wu, Ze Wang, Seonglae Cho, Yufei Yang, Adriano Koshiyama · May 8, 2026 · Citations: 0
When a tool-calling agent picks the wrong tool, the failure is invisible until execution: the email gets sent, the meeting gets missed.
Qingyu Lu, Liang Ding, Kanjian Zhang, Jinxia Zhang, Dacheng Tao · Jan 19, 2026 · Citations: 0
In this work, we present a comprehensive evaluation of dLLMs (e.g., LLaDA, Dream) across two distinct agentic paradigms: Embodied Agents (requiring long-horizon planning) and Tool-Calling Agents (requiring precise formatting).
Yutao Hou, Zeguan Xiao, Fei Yu, Yihan Jiang, Ma Shuguang · Jun 5, 2025 · Citations: 0
Existing robustness evaluations predominantly rely on hand-crafted templates or a limited set of perturbation rules.
Chenxi Wang, Zhuoyun Yu, Xin Xie, Wuguannan Yao, Runnan Fang · Apr 6, 2026 · Citations: 0
Learning from experience is critical for building capable large language model (LLM) agents, yet prevailing self-evolving paradigms remain inefficient: agents learn in isolation, repeatedly rediscover similar behaviors from limited…
Shangqing Tu, Yaxuan Li, Yushi Bai, Lei Hou, Juanzi Li · Oct 9, 2025 · Citations: 0
Our method features a specialized judge model trained with out-of-distribution data (AIME 2022, AIME 2023, and MATH-500) using oversampling techniques to accurately predict answer equivalence from partial reasoning traces, achieving 0.7072…
Junhao Su, Yuanliang Wan, Junwei Yang, Hengyu Shi, Tianyang Han · Sep 23, 2025 · Citations: 0
The agent produces a short yet precise reflection: it diagnoses the failure using evidence from the previous step and then proposes a correct, executable follow-up call.
Teng Wang, Zhangyi Jiang, Zhenqi He, Shenyang Tong, Wenhan Yang · Mar 16, 2025 · Citations: 0
Empirical results on the PRM800K dataset show that HRM, together with HNC, provides more stable and reliable evaluations than PRM.
Xuan Qi · Apr 2, 2026 · Citations: 0
Chain-of-thought (CoT) reasoning is widely assumed to improve agent performance, but the relationship between reasoning length and accuracy in structured tool-use settings remains poorly understood.
Haomin Zhuang, Hojun Yoo, Xiaonan Luo, Kehan Guo, Xiangliang Zhang · Apr 2, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Cai Zhou, Zekai Wang, Menghua Wu, Qianyu Julie Zhu, Flora C. Shi · Apr 1, 2026 · Citations: 0
Under zero-shot out-of-domain settings, it improves MATH-500 savings from 24.8% of the static calibration baseline to 67.0% while maintaining a low empirical error rate, and the same trend holds across model families and downstream…
Jack Young · Apr 1, 2026 · Citations: 0
Using roughly 48 execution-verified HumanEval training solutions, tuning a single initial state matrix per recurrent layer, with zero inference overhead, outperforms LoRA by +10.8 pp (p < 0.001) on HumanEval.
Liancheng Fang, Aiwei Liu, Henry Peng Zou, Yankai Chen, Enze Ma · Apr 1, 2026 · Citations: 0
Experiments across a range of reasoning benchmarks including MATH500, AIME24/25, HumanEval, and MBPP show that our approach yields better exploration-quality tradeoff than both random and low-confidence remasking.
Mohamad Zbib, Mohamad Bazzi, Ammar Mohanna, Hasan Abed Al Kader Hammoud, Bernard Ghanem · Mar 27, 2026 · Citations: 0
Measured by acceptance length, task-specific training yields clear specialization: MathInstruct-trained drafts are strongest on reasoning benchmarks, while ShareGPT-trained drafts are strongest on MT-Bench.
Heecheol Yun, Kwangmin Ki, Junghyun Lee, Eunho Yang · Oct 17, 2025 · Citations: 0
Our experiments on diverse benchmarks, including MATH500 and BBH, demonstrate that SAFE outperforms existing methods in both accuracy and efficiency, with gains achieved even when ensembling fewer than 1% of tokens.
Kunfeng Chen, Qihuang Zhong, Juhua Liu, Bo Du, Dacheng Tao · Mar 12, 2026 · Citations: 0
Tool-DC (TF) brings up to +25.10% average gains against the baseline on BFCL and ACEBench benchmarks, while Tool-DC (TB) enables Qwen2.5-7B to achieve comparable or even better performance than proprietary LLMs, e.g., OpenAI o3 and…
Konrad Staniszewski, Adrian Łańcucki · Nov 3, 2025 · Citations: 0
We test KVTC with Llama 3, Mistral NeMo, and R1-Qwen 2.5 models across benchmarks including AIME25, GSM8K, LiveCodeBench, LongBench, MATH-500, MMLU, Qasper and RULER.
Bo Jiang · Mar 8, 2026 · Citations: 0
We introduce a taxonomy of three defense categories -- output perturbation, data poisoning, and information throttling -- and evaluate nine defense configurations using a standardized pipeline with Qwen3-14B as teacher and…
Shaoxiong Zhan, Yanlin Lai, Ziyu Lu, Dahua Lin, Ziqing Yang · Aug 7, 2025 · Citations: 0
Existing synthesis methods largely rely on transforming human-written templates, limiting both diversity and scalability.
Chongyu Fan, Yihua Zhang, Jinghan Jia, Alfred Hero, Sijia Liu · Jun 4, 2025 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Dongxu Zhang, Yujun Wu, Yiding Sun, Jinnan Yang, Ning Yang · Aug 7, 2025 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Ngoc Bui, Shubham Sharma, Simran Lamba, Saumitra Mishra, Rex Ying · Dec 3, 2025 · Citations: 0
Across mathematical reasoning (GSM8K, MATH-500, AIME24), procedural generation (LongProc), conversational long-memory benchmarks (LongMemEval), and long-context understanding (LongBenchV2 and SCBench), TRIM-KV consistently outperforms…
Yuchen Yan, Yongliang Shen, Yang Liu, Jin Jiang, Mengdi Zhang · Mar 9, 2025 · Citations: 0
Experiments across multiple model architectures demonstrate that our approach reduces computational costs while improving performance, with Qwen2.5-Math-7B showing 3-11% improvements across MATH500, AIME24, and GPQA_diamond benchmarks.
Rulin Shao, Shuyue Stella Li, Rui Xin, Scott Geng, Yiping Wang · Jun 12, 2025 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Yuanda Xu, Hejian Sang, Zhengze Zhou, Ran He, Zhipeng Wang · Feb 24, 2026 · Citations: 0
Evaluated on MATH-500 and AIME 2025, ACE composes seamlessly with existing methods and consistently improves the full Pass@k spectrum across all three model families and benchmarks.
Ahsan Bilal, Ahmed Mohsin, Muhammad Umer, Ali Subhan, Hassan Rizwan · Feb 1, 2026 · Citations: 0
For each problem, the agent runs multiple inference iterations.
Mert Cemri, Nived Rajaraman, Rishabh Tiwari, Xiaoxuan Liu, Kurt Keutzer · Jun 15, 2025 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Buze Zhang, Jinkai Tao, Zilang Zeng, Neil He, Ali Maatouk · Feb 16, 2026 · Citations: 0
Our experiments across diverse benchmarks demonstrate that MoSLoRA consistently outperforms strong baselines, achieving up to 5.6% improvement on MATH500 and 15.9% on MAWPS.