HFEPX Benchmark Hub
AIME or BFCL or MMLU Benchmark Papers
Updated from current HFEPX corpus (2026-09-13). This page tracks 60 papers reporting AIME or BFCL or MMLU benchmark evidence, with protocol and metric context for comparison.
HFEPX Benchmark Hub
Updated from current HFEPX corpus (2026-09-13). This page tracks 60 papers reporting AIME or BFCL or MMLU benchmark evidence, with protocol and metric context for comparison.
Use this page for benchmark-matched method comparisons and eval protocol selection. Quality band: High .
High-Signal Coverage
100.0%
60 / 60 sampled papers are not low-signal flagged.
Replication-Ready Set
40
Papers with explicit benchmark + metric + eval mode fields.
Quality Controls
5.0%
3 papers report calibration/adjudication/IAA controls.
Primary action: Start with the top 2 benchmark-matched papers, then compare evaluation modes in the protocol matrix.
Ranked by protocol completeness so you can quickly find papers suitable for comparison studies.
Aug 19, 2026 · Citations: 0 · Score: 8.5
Eval: Automatic Metrics · Metrics: Accuracy
Aug 27, 2026 · Citations: 0 · Score: 8.5
Eval: Automatic Metrics · Metrics: Spearman
Aug 21, 2026 · Citations: 0 · Score: 8.5
Eval: Automatic Metrics · Metrics: Accuracy
Jul 2, 2026 · Citations: 0 · Score: 8.0
Eval: Automatic Metrics, Simulation Env · Metrics: Accuracy
Jun 24, 2026 · Citations: 0 · Score: 8.0
Eval: Automatic Metrics · Metrics: Accuracy
May 30, 2026 · Citations: 0 · Score: 7.5
Eval: Simulation Env · Metrics: Precision
Compare protocol ingredients quickly before deep-reading full papers.
| Paper | Eval Modes | Human Feedback | Metrics | Quality Controls |
|---|---|---|---|---|
| Decomposing Wrong-Consensus Agreement in LLM Self-Consistency Aug 19, 2026 | Automatic Metrics | Pairwise Preference | Accuracy | Not reported |
| TwinKV: A Composable Repair Pass for KV Cache Eviction via Pairwise Key Redundancy Aug 27, 2026 | Automatic Metrics | Pairwise Preference | Spearman | Not reported |
| Memory Augmentation Unlocks Efficient Chain-of-Thought Reasoning Aug 21, 2026 | Automatic Metrics | Demonstrations | Accuracy, Latency | Not reported |
| Will Scaling Improve Social Simulation with LLMs? Jul 2, 2026 | Automatic Metrics, Simulation Env | Not reported | Accuracy | Calibration |
| Cliff Tokens: Identifying Single-Token Failure Triggers in LLM Mathematical Reasoning Jun 24, 2026 | Automatic Metrics | Pairwise Preference | Accuracy, Pass@64 | Not reported |
| Skill or Skip? Learning Selective Skill Invocation in Agentic Tasks via Dual-Granularity Preference Learning May 30, 2026 | Simulation Env | Pairwise Preference | Precision, Task success | Not reported |
| Breaking MCP with Function Hijacking Attacks: Novel Threats for Function Calling and Agentic Models Apr 22, 2026 | Automatic Metrics | Pairwise Preference, Red Team | Jailbreak success rate | Not reported |
| Hidden Measurement Error in LLM Pipelines Distorts Annotation, Evaluation, and Benchmarking Apr 13, 2026 | Llm As Judge | Demonstrations | Precision, Agreement | Not reported |
| Diagnosing Translated Benchmarks: An Automated Quality Assurance Study of the EU20 Benchmark Suite Apr 2, 2026 | Automatic Metrics | Not reported | Accuracy | Calibration, Gold Questions |
| PubMed Reasoner: Dynamic Reasoning-based Retrieval for Evidence-Grounded Biomedical Question Answering Mar 28, 2026 | Llm As Judge, Automatic Metrics | Expert Verification | Accuracy, Relevance | Not reported |
Gap: Human feedback
Human feedback is present in 11 of 60 papers.
Gap: Quality controls
Quality controls is present in 3 of 60 papers.
Strong: Benchmarks
Benchmarks is present in 60 of 60 papers.
Strong: Metrics
Metrics is present in 43 of 60 papers.
Gap: Known rater population
Known rater population is present in 8 of 60 papers.
Gap: Known annotation unit
Known annotation unit is present in 10 of 60 papers.
Evaluation Modes
Human Feedback Mix
Top Benchmarks
Top Metrics
Chishui Chen, Jiaye Lin, Te Sun, Yi Yang, Junxi Wang · May 30, 2026 · Citations: 0
Agent skills are callable procedural modules that provide reusable knowledge and execution policies for complex agentic tasks.
Lizhuo Zhang, Mengmeng Tang, Chenfeng Long, Xiaoyong Tang, Xiang Luo · Aug 19, 2026 · Citations: 0
The mechanical reference is leak-free: each case's preference and accuracy are estimated from its other runs only.
Bo Liu, Simon Yu, Yiding Jiang, Ao Qu, Andrew Zhao · Aug 19, 2026 · Citations: 0
For language agents, existing training environment pools (hand-curated, statically synthesized, or frozen-verifier) keep the goal distribution fixed as the learner scales.
Yannis Belkhiter, Giulio Zizzo, Sergio Maffeis, Seshu Tirupathi, John D. Kelleher · Apr 22, 2026 · Citations: 0
The growth of agentic AI has drawn significant attention to function calling Large Language Models (LLMs), which are designed to extend the capabilities of AI-powered system by invoking external functions.
Minsoo Song, Chanjun Park · Aug 31, 2026 · Citations: 0
To provide a scalable diagnostic approach, we propose a two-component probabilistic framework for auditing MCQA benchmarks using model output distributions.
Dhruv Deshmukh, Saurabh Goyal, Nipun Kwatra, Ramachandran Ramjee · Dec 18, 2025 · Citations: 0
Kascade achieves up to 4.1x speedup in decode attention and 2.2x speedup in prefill attention over FlashAttention-3 baseline on H100 GPUs while closely matching dense attention accuracy on long-context benchmarks such as LongBench and…
Hong Chen, Yudong Zeng, Yongwei Huang, Zuhao Ouyang, Dongnan Zheng · Aug 27, 2026 · Citations: 0
We introduce TwinKV, a training-free, attention-free redundancy signal that detects whether a token's key has a near-duplicate elsewhere in context.
Ruotong Liao, Nikolai Röhrich, Xiaohan Wang, Yuhui Zhang, Yasaman Samadzadeh · Mar 2, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Haodong Zhao, Jidong Li, Zhaomin Wu, Tianjie Ju, Zhuosheng Zhang · Sep 25, 2025 · Citations: 0
Understanding persuasion is critical for the safety and reliability of multi-agent systems built on large language models (LLMs).
Simeng Zhang, Yilong Chen, Wenyuan Zhang, Zhenyu Zhang, Yao Chen · Aug 21, 2026 · Citations: 0
Based on this principle, we propose Memory-Augmented Compression, a training-free framework that constructs reusable reasoning memories from historical traces and retrieves them as prefill-side scaffolds.
Zhibo Zhang, Zhen Ouyang, Ling Shi, Kailong Wang · Aug 21, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Qian Kou, Xiaofeng Shi, Xiaosong Qiu, Hua Zhou · Aug 20, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Zhixuan Liu, Zhichen Dong, Yuanfu Wang, Chao Yang · Aug 13, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Shailja Thakur, Sungeun An, Chad DeLuca, Hima Patel · Aug 12, 2026 · Citations: 0
A benchmark score comes from a single phrasing of each problem.
Dylan Bouchard, Mohit Singh Chauhan · Aug 12, 2026 · Citations: 0
For LLM agents, however, the unit of observation is an interactive trajectory, where the model can ask clarifying questions, call tools, update state, and make intermediate decisions whose errors propagate to the final outcome.
Zhenyu Zhao, Sander Land, Daniel M. Bikel, Waseem Alshikh · Apr 29, 2026 · Citations: 0
Across three model families and five mathematical reasoning benchmarks, our approach shortens reasoning traces by 8.1% on average; under a TOST equivalence analysis at a +/- 2pp margin, accuracy is equivalent or inconclusive on 13/15 model…
Caleb Ziems, William Held, Su Doga Karaca, David Grusky, Tatsunori Hashimoto · Jul 2, 2026 · Citations: 0
We use scaling laws to study the relationship between LLMs' compute scale, general capability benchmarks, and the fidelity of social simulation in three representative sub-domains: opinion modeling, behavioral simulation, and longitudinal…
Xiangdong Zhang, Debing Zhang, Shaofeng Zhang, Xiaohan Qin, Yu Cheng · May 24, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Andikawati P Widjaja, Yongjun Kim, Hyounghun Kim, Jaeho Lee · Jul 2, 2026 · Citations: 0
Across eight benchmarks (including MMLU, GSM8K, and RULER) and three model families (Qwen2.5, Llama3.2, Gemma4), PartRep retains most of the gains of full repetition while using only 59.4\% of its KV cache and 79.0\% of its prefill FLOPs.
Kevin Yandoka Denamganaï · May 27, 2026 · Citations: 0
Self-evolving scientific agents capable of conquering the hard tail of formal mathematics require Compositional Learning Behaviours (CLBs) -- the capacity to ground and recombine novel symbolic structures in context, beyond mere…
Subhadip Mitra · Jun 28, 2026 · Citations: 0
Inference-time safety methods for large language models have proliferated, yet no systematic comparison exists.
Steven Kolawole, Virginia Smith · Jun 25, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Xiangyue Liu, Zijian Zhang, Miles Yang, Zhao Zhong, Liefeng Bo · Apr 9, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Jaeyong Ko, Pilsung Kang, Yukyung Lee · Jun 24, 2026 · Citations: 0
Across seven models and three mathematical reasoning benchmarks (GSM1K, MATH500, AIME 2025), cliff tokens act as failure triggers; deleting the first cliff token and resampling recovers pass@64 to 1.0, while keeping it limits recovery to…
Tianyu Dong, Yangyang Liu, Jiang Zhou, Xinwei Wu, Xiaohu Zhao · Jun 24, 2026 · Citations: 0
We conduct experiments on 2 LLMs across 5 low-resource languages and 3 benchmarks.
Chenhao Dang, Jing Ma, Mingjie Liao · Jun 23, 2026 · Citations: 0
On The Pile benchmark, HDS reaches the final validation perplexity of the next best method with 44% fewer training iterations.
Fengfeng Liang, Yuechen Zhang, Jiaya Jia · Jun 23, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Ali Asgarov, Umid Suleymanov, Aadyant Khatri · Oct 31, 2025 · Citations: 0
We introduce SIGMA (Search-Augmented On-Demand Knowledge Integration for AGentic Mathematical reAsoning), a unified framework that orchestrates specialized agents to independently reason, perform targeted searches, and synthesize findings…
Glenn Matlin, Chandreyi Chakraborty, Saehee Eom, Mika Okamoto, Rayan Castilla · Jun 17, 2026 · Citations: 0
Training-data attribution measures how strongly each training document influences a model's predictions on a benchmark, but document-level scores are too noisy to identify which corpus regions support which capabilities, and prior work has…
Anany Kotawala · May 28, 2026 · Citations: 0
Across two public LLM leaderboards, many displayed pairwise rankings do not meet a conventional paired-test resolution target under the actual paired evaluation design: 11 of 40 Open LLM Leaderboard v1 pairwise comparisons and 4 of 9…
Siddharth Boppana, Annabel Ma, Max Loeffler, Raphael Sarfati, Eric Bigelow · Mar 5, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Tanmoy Chakraborty, Ayan Sengupta, Suparna Bhattacharya, Partha Pratim Chakrabarti, Amlan Chakrabarti · May 28, 2026 · Citations: 0
Large language models (LLMs) frequently achieve impressive scores on standardized benchmarks, yet accuracy alone offers a limited view of their capabilities.
Woomin Song, Saket Dingliwal, Sai Muralidhar Jayanthi, Bhavana Ganesh, Jinwoo Shin · Jun 5, 2025 · Citations: 0
Extensive evaluations across multiple models and reasoning tasks (AIME-2024, GPQA-Diamond, and LiveCodeBench) demonstrate that STAND reduces inference latency by 60-65% compared to standard autoregressive decoding while maintaining…
Andrea Sassella, Andrea Chizzola, Tommaso Bianchi, Luca Alessandrelli, Mark James Carman · May 8, 2026 · Citations: 0
This report benchmarks the performance of ENGINEERING Ingegneria Informatica S.p.A.'s EngGPT2MoE-16B-A3B LLM, a 16B parameter Mixture of Experts (MoE) model with 3B active parameters.
Zekun Wu, Ze Wang, Seonglae Cho, Yufei Yang, Adriano Koshiyama · May 8, 2026 · Citations: 0
When a tool-calling agent picks the wrong tool, the failure is invisible until execution: the email gets sent, the meeting gets missed.
Qiyong Zhong, Mao Zheng, Mingyang Song, Xin Lin, Jie Sun · May 8, 2026 · Citations: 0
To address this, we propose SOD, a step-wise on-policy distillation framework for small language model agents, which adaptively reweights distillation strength at each step based on step-level divergence.
Tsuyoshi Okita · May 8, 2026 · Citations: 0
With a 7B-parameter LLM whose weights are entirely frozen, CIKA achieves 69.7\% on the contamination-free Omni-MATH-Rule benchmark and 64.0\% overall, compared to 60.5\% for o1-mini, and 97.2\% on GSM8K, 46--50\% on AIME 2024--2026, and…
Qianjia Cheng, Yuchen Zhang, Zhilin Wang, Yuxin Zuo, Shunkai Zhang · May 7, 2026 · Citations: 0
Paradoxically, we observe that tool-enabled evaluation can degrade reasoning performance even when the strong thinking models make almost no actual tool calls.
Solomon Messing · Apr 13, 2026 · Citations: 0
LLM evaluations drive which models get deployed, what safety standards get adopted, which research conclusions get published, and how projections of AI's labor-market impact get made.
Xiaoyu Xu, Minxin Du, Kun Fang, Yaxin Xiao, Zhicong Huang · Jan 29, 2026 · Citations: 0
Furthermore, to facilitate rigorous evaluation, we introduce PCH, a unified benchmark encompassing Personal, Copyrighted, and Harmful content, alongside two symmetric metrics, Forget Degree (F.D.) and Retain Utility (R.U.), to…
Weiqin Wang, Yile Wang, Kehao Chen, Hui Huang · Dec 17, 2025 · Citations: 0
We conduct experiments across various models and benchmarks, experimental results show that SCOPE consistently outperforms recent baselines.
Pere Martra · Dec 27, 2025 · Citations: 0
We evaluated seven expansion ratio configurations using comprehensive benchmarks assessing factual knowledge, mathematical reasoning, language comprehension, instruction-following, and truthfulness.
Weitao Li, Boran Xiang, Xiaolong Wang, Zhinan Gou, Weizhi Ma · Aug 8, 2025 · Citations: 0
Experiments on open-domain QA, MMLU-Pro, medical, and mathematical reasoning tasks show that UR^2, built on Qwen-2.5-3/7B and LLaMA-3.1-8B, consistently outperforms existing RAG and RL baselines, and achieves performance comparable to…
Qingyu Lu, Liang Ding, Kanjian Zhang, Jinxia Zhang, Dacheng Tao · Jan 19, 2026 · Citations: 0
In this work, we present a comprehensive evaluation of dLLMs (e.g., LLaDA, Dream) across two distinct agentic paradigms: Embodied Agents (requiring long-horizon planning) and Tool-Calling Agents (requiring precise formatting).
Yijia Fan, Jusheng Zhang, Kaitong Cai, Jing Yang, Chengpei Tang · Nov 17, 2025 · Citations: 0
To address this, we introduce the Dynamic Auction-based Language Agent (DALA), a novel framework that treats communication bandwidth as a scarce and tradable resource.
Chenxi Wang, Zhuoyun Yu, Xin Xie, Wuguannan Yao, Runnan Fang · Apr 6, 2026 · Citations: 0
Learning from experience is critical for building capable large language model (LLM) agents, yet prevailing self-evolving paradigms remain inefficient: agents learn in isolation, repeatedly rediscover similar behaviors from limited…
Shangqing Tu, Yaxuan Li, Yushi Bai, Lei Hou, Juanzi Li · Oct 9, 2025 · Citations: 0
Our method features a specialized judge model trained with out-of-distribution data (AIME 2022, AIME 2023, and MATH 500) using oversampling techniques to accurately predict answer equivalence from partial reasoning traces, achieving 0.7072…
Junhao Su, Yuanliang Wan, Junwei Yang, Hengyu Shi, Tianyang Han · Sep 23, 2025 · Citations: 0
The agent produces a short yet precise reflection: it diagnoses the failure using evidence from the previous step and then proposes a correct, executable follow-up call.
Pengxiang Li, Yefan Zhou, Dilxat Muhtar, Lu Yin, Shilin Yan · Aug 27, 2025 · Citations: 0
Empirical evaluations of LLaDA-8B and Dream-7B across multiple tasks show that Prophet reduces the number of decoding steps by up to 3.4x while preserving high generation quality.
Zhichao Wang · Oct 27, 2025 · Citations: 0
This paper proposes Group-relative Implicit Fine-Tuning (GIFT), a reinforcement learning framework for aligning large language models (LLMs) that unifies on-policy optimization with implicit preference learning.
Xuan Qi · Apr 2, 2026 · Citations: 0
Chain-of-thought (CoT) reasoning is widely assumed to improve agent performance, but the relationship between reasoning length and accuracy in structured tool-use settings remains poorly understood.
Klaudia Thellmann, Bernhard Stadler, Michael Färber · Apr 2, 2026 · Citations: 0
Machine-translated benchmark datasets reduce costs and offer scale, but noise, loss of structure, and uneven quality weaken confidence.
Ante Wang, Weizhi Ma, Yang Liu · Nov 18, 2025 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Zhenpeng Su, Leiyu Pan, Xue Bai, Dening Liu, Guanting Dong · Aug 11, 2025 · Citations: 0
We present Klear-Reasoner, a model with long reasoning capabilities that demonstrates careful deliberation during problem solving, achieving outstanding performance across multiple benchmarks.
M. Ali Bayram, Ali Arda Fincan, Ahmet Semih Gümüş, Sercan Karakaş, Banu Diri · Aug 19, 2025 · Citations: 0
We further validate practical utility with downstream sentence embedding benchmarks under a strict random initialization control to isolate tokenizer inductive bias.
G. Ciarfaglia, A. Rosanova, S. Cipolla, J. Bartoli, A. Di Domenico · Mar 17, 2026 · Citations: 0
EngGPT2 is trained on 2.5 trillion tokens - less than Qwen3's 36T or Llama3's 15T - and delivers performance on key benchmarks, including MMLU-Pro, GSM8K, IFEval and HumanEval, comparable to dense models in the 8B-16B range, while requiring…
Yiqing Zhang, Xiaozhong Liu, Fabricio Murai · Mar 28, 2026 · Citations: 0
In this context, we introduce PubMed Reasoner, a biomedical QA agent composed of three stages: self-critic query refinement evaluates MeSH terms for coverage, alignment, and redundancy to enhance PubMed queries based on partial (metadata)…
Zara Siddique, Irtaza Khalid, Liam D. Turner, Luis Espinosa-Anke · Mar 7, 2025 · Citations: 0
When optimized on the BBQ dataset, our individually tuned steering vectors achieve average improvements of 12.8% on BBQ, 8.3% on CLEAR-Bias, and 1% on StereoSet, and show improvements over prompting and Self-Debias in all cases, and…
Richard J. Young · Mar 27, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Hao Liang, Zhengyang Zhao, Meiyi Qiang, Mingrui Chen, Lu Ma · Mar 27, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.