HFEPX Benchmark Hub
BFCL or HumanEval+ Benchmark Papers
Updated from current HFEPX corpus (2026-09-13). This page tracks 11 papers reporting BFCL or HumanEval+ benchmark evidence, with protocol and metric context for comparison.
HFEPX Benchmark Hub
Updated from current HFEPX corpus (2026-09-13). This page tracks 11 papers reporting BFCL or HumanEval+ benchmark evidence, with protocol and metric context for comparison.
Use this page for benchmark-matched method comparisons and eval protocol selection. Quality band: Developing .
High-Signal Coverage
100.0%
11 / 11 sampled papers are not low-signal flagged.
Replication-Ready Set
8
Papers with explicit benchmark + metric + eval mode fields.
Quality Controls
0.0%
0 papers report calibration/adjudication/IAA controls.
Primary action: Start with the top 2 benchmark-matched papers, then compare evaluation modes in the protocol matrix.
Ranked by protocol completeness so you can quickly find papers suitable for comparison studies.
May 30, 2026 · Citations: 0 · Score: 7.5
Eval: Simulation Env · Metrics: Precision
Apr 22, 2026 · Citations: 0 · Score: 7.5
Eval: Automatic Metrics · Metrics: Jailbreak success rate
Aug 12, 2026 · Citations: 0 · Score: 6.5
Eval: Automatic Metrics · Metrics: Cost
May 8, 2026 · Citations: 0 · Score: 6.0
Eval: Automatic Metrics · Metrics: Accuracy
Apr 6, 2026 · Citations: 0 · Score: 6.0
Eval: Automatic Metrics · Metrics: Task success
Apr 2, 2026 · Citations: 0 · Score: 6.0
Eval: Automatic Metrics · Metrics: Accuracy
Compare protocol ingredients quickly before deep-reading full papers.
Gap: Human feedback
Human feedback is present in 2 of 11 papers.
Gap: Quality controls
Quality controls is present in 0 of 11 papers.
Strong: Benchmarks
Benchmarks is present in 11 of 11 papers.
Strong: Metrics
Metrics is present in 8 of 11 papers.
Gap: Known rater population
Known rater population is present in 1 of 11 papers.
Moderate: Known annotation unit
Known annotation unit is present in 3 of 11 papers.
Evaluation Modes
Human Feedback Mix
Top Benchmarks
Top Metrics
Chishui Chen, Jiaye Lin, Te Sun, Yi Yang, Junxi Wang · May 30, 2026 · Citations: 0
Agent skills are callable procedural modules that provide reusable knowledge and execution policies for complex agentic tasks.
Bo Liu, Simon Yu, Yiding Jiang, Ao Qu, Andrew Zhao · Aug 19, 2026 · Citations: 0
For language agents, existing training environment pools (hand-curated, statically synthesized, or frozen-verifier) keep the goal distribution fixed as the learner scales.
Yannis Belkhiter, Giulio Zizzo, Sergio Maffeis, Seshu Tirupathi, John D. Kelleher · Apr 22, 2026 · Citations: 0
The growth of agentic AI has drawn significant attention to function calling Large Language Models (LLMs), which are designed to extend the capabilities of AI-powered system by invoking external functions.
Dylan Bouchard, Mohit Singh Chauhan · Aug 12, 2026 · Citations: 0
For LLM agents, however, the unit of observation is an interactive trajectory, where the model can ask clarifying questions, call tools, update state, and make intermediate decisions whose errors propagate to the final outcome.
Andrea Sassella, Andrea Chizzola, Tommaso Bianchi, Luca Alessandrelli, Mark James Carman · May 8, 2026 · Citations: 0
This report benchmarks the performance of ENGINEERING Ingegneria Informatica S.p.A.'s EngGPT2MoE-16B-A3B LLM, a 16B parameter Mixture of Experts (MoE) model with 3B active parameters.
Zekun Wu, Ze Wang, Seonglae Cho, Yufei Yang, Adriano Koshiyama · May 8, 2026 · Citations: 0
When a tool-calling agent picks the wrong tool, the failure is invisible until execution: the email gets sent, the meeting gets missed.
Qingyu Lu, Liang Ding, Kanjian Zhang, Jinxia Zhang, Dacheng Tao · Jan 19, 2026 · Citations: 0
In this work, we present a comprehensive evaluation of dLLMs (e.g., LLaDA, Dream) across two distinct agentic paradigms: Embodied Agents (requiring long-horizon planning) and Tool-Calling Agents (requiring precise formatting).
Chenxi Wang, Zhuoyun Yu, Xin Xie, Wuguannan Yao, Runnan Fang · Apr 6, 2026 · Citations: 0
Learning from experience is critical for building capable large language model (LLM) agents, yet prevailing self-evolving paradigms remain inefficient: agents learn in isolation, repeatedly rediscover similar behaviors from limited…
Junhao Su, Yuanliang Wan, Junwei Yang, Hengyu Shi, Tianyang Han · Sep 23, 2025 · Citations: 0
The agent produces a short yet precise reflection: it diagnoses the failure using evidence from the previous step and then proposes a correct, executable follow-up call.
Xuan Qi · Apr 2, 2026 · Citations: 0
Chain-of-thought (CoT) reasoning is widely assumed to improve agent performance, but the relationship between reasoning length and accuracy in structured tool-use settings remains poorly understood.
Kunfeng Chen, Qihuang Zhong, Juhua Liu, Bo Du, Dacheng Tao · Mar 12, 2026 · Citations: 0
Tool-DC (TF) brings up to +25.10% average gains against the baseline on BFCL and ACEBench benchmarks, while Tool-DC (TB) enables Qwen2.5-7B to achieve comparable or even better performance than proprietary LLMs, e.g., OpenAI o3 and…