Skip to content
OpenTrain AIFor AI Companies
← Back to explorer

HFEPX Benchmark Hub

BFCL or HumanEval+ Benchmark Papers

Updated from current HFEPX corpus (2026-09-13). This page tracks 11 papers reporting BFCL or HumanEval+ benchmark evidence, with protocol and metric context for comparison.

Papers: 11 Last published: Aug 19, 2026 Global RSS

Researcher Quick Triage

Use this page for benchmark-matched method comparisons and eval protocol selection. Quality band: Developing .

High-Signal Coverage

100.0%

11 / 11 sampled papers are not low-signal flagged.

Replication-Ready Set

8

Papers with explicit benchmark + metric + eval mode fields.

Quality Controls

0.0%

0 papers report calibration/adjudication/IAA controls.

  • 11 papers explicitly name benchmark datasets in the sampled set.
  • 8 papers report at least one metric term in metadata extraction.
  • Start with the ranked shortlist below before reading all papers.

Primary action: Start with the top 2 benchmark-matched papers, then compare evaluation modes in the protocol matrix.

Why This Matters (Expanded)

Why This Matters For Eval Research

  • Use this page to compare BFCL or HumanEval+ papers by evaluation mode, metric, and evidence quality before reusing reported results.
Protocol Notes (Expanded)

Protocol Takeaways

  • BFCL or HumanEval+ papers are often paired with automatic_metrics, simulation_env.

Benchmark Interpretation

  • BFCL: 11 papers
  • Acebench: 2 papers
  • ALFWorld: 1 papers
  • ARC-Challenge: 1 papers

Metric Interpretation

  • accuracy: 3 papers
  • precision: 2 papers
  • task success: 2 papers
  • cost: 1 papers

Start Here (Benchmark-Matched First 6)

Ranked by protocol completeness so you can quickly find papers suitable for comparison studies.

Protocol Matrix (Top 10)

Compare protocol ingredients quickly before deep-reading full papers.

Paper Eval Modes Human Feedback Metrics Quality Controls
Skill or Skip? Learning Selective Skill Invocation in Agentic Tasks via Dual-Granularity Preference Learning

May 30, 2026

Simulation Env Pairwise Preference Precision, Task success Not reported
Breaking MCP with Function Hijacking Attacks: Novel Threats for Function Calling and Agentic Models

Apr 22, 2026

Automatic Metrics Pairwise Preference, Red Team Jailbreak success rate Not reported
Beyond Single-Turn Confidence: Trajectory-Adapted Uncertainty Quantification for LLM Agents

Aug 12, 2026

Automatic Metrics Not reported Cost Not reported
Tool Calling is Linearly Readable and Steerable in Language Models

May 8, 2026

Automatic Metrics Not reported Accuracy Not reported
SkillX: Automatically Constructing Skill Knowledge Bases for Agents

Apr 6, 2026

Automatic Metrics Not reported Task success Not reported
Brief Is Better: Non-Monotonic Chain-of-Thought Budget Effects in Function-Calling Language Agents

Apr 2, 2026

Automatic Metrics Not reported Accuracy Not reported
The Bitter Lesson of Diffusion Language Models for Agentic Workflows: A Comprehensive Reality Check

Jan 19, 2026

Automatic Metrics Not reported Precision, Latency Not reported
Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions

Sep 23, 2025

Automatic Metrics Not reported Accuracy Not reported
SPADE: Self-Play in Adaptive Synthetic Executable Environments

Aug 19, 2026

Not reported Not reported Not reported Not reported
Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs

May 8, 2026

Not reported Not reported Not reported Not reported
Researcher Workflow (Detailed)

Checklist

  • Gap: Human feedback

    Human feedback is present in 2 of 11 papers.

  • Gap: Quality controls

    Quality controls is present in 0 of 11 papers.

  • Strong: Benchmarks

    Benchmarks is present in 11 of 11 papers.

  • Strong: Metrics

    Metrics is present in 8 of 11 papers.

  • Gap: Known rater population

    Known rater population is present in 1 of 11 papers.

  • Moderate: Known annotation unit

    Known annotation unit is present in 3 of 11 papers.

Strengths

  • Benchmarks is present in 11 of 11 papers.
  • Metrics is present in 8 of 11 papers.
  • Agentic evaluation is present in 8 of 11 papers.

Known Gaps

  • Human feedback is present in 2 of 11 papers.
  • Quality controls is present in 0 of 11 papers.
  • Known rater population is present in 1 of 11 papers.

Suggested Next Analyses

  • Review the most recent BFCL or HumanEval+ papers first, then compare reported metrics and quality-control context before treating results as comparable.

Recommended Queries

Known Limitations
  • This synthetic persisted page is generated from extraction data because the cached benchmark payload was missing for either-bfcl-or-humaneval.
Research Utility Snapshot (Detailed)

Evaluation Modes

  • Automatic Metrics (7)
  • Simulation Env (1)

Human Feedback Mix

  • None (9)
  • Pairwise Preference (2)
  • Red Team (1)

Top Benchmarks

  • BFCL (11)
  • Acebench (2)
  • ALFWorld (1)
  • ARC-Challenge (1)

Top Metrics

  • Accuracy (3)
  • Precision (2)
  • Task success (2)
  • Cost (1)

Top Papers On This Benchmark

Related Benchmark Hubs