HFEPX Benchmark Hub
AIME or GSM8K or MMLU Benchmark Papers
Updated from current HFEPX corpus (2026-09-29). This page tracks 60 papers reporting AIME or GSM8K or MMLU benchmark evidence, with protocol and metric context for comparison.
HFEPX Benchmark Hub
Updated from current HFEPX corpus (2026-09-29). This page tracks 60 papers reporting AIME or GSM8K or MMLU benchmark evidence, with protocol and metric context for comparison.
Use this page for benchmark-matched method comparisons and eval protocol selection. Quality band: High .
High-Signal Coverage
100.0%
60 / 60 sampled papers are not low-signal flagged.
Replication-Ready Set
37
Papers with explicit benchmark + metric + eval mode fields.
Quality Controls
3.3%
2 papers report calibration/adjudication/IAA controls.
Primary action: Start with the top 2 benchmark-matched papers, then compare evaluation modes in the protocol matrix.
Ranked by protocol completeness so you can quickly find papers suitable for comparison studies.
Aug 19, 2026 · Citations: 0 · Score: 8.0
Eval: Automatic Metrics · Metrics: Accuracy
Aug 27, 2026 · Citations: 0 · Score: 8.0
Eval: Automatic Metrics · Metrics: Spearman
Aug 21, 2026 · Citations: 0 · Score: 8.0
Eval: Automatic Metrics · Metrics: Accuracy
Jul 2, 2026 · Citations: 0 · Score: 8.0
Eval: Automatic Metrics, Simulation Env · Metrics: Accuracy
Jun 24, 2026 · Citations: 0 · Score: 7.5
Eval: Automatic Metrics · Metrics: Accuracy
May 3, 2026 · Citations: 0 · Score: 7.5
Eval: Automatic Metrics · Metrics: Accuracy
Compare protocol ingredients quickly before deep-reading full papers.
| Paper | Eval Modes | Human Feedback | Metrics | Quality Controls |
|---|---|---|---|---|
| Decomposing Wrong-Consensus Agreement in LLM Self-Consistency Aug 19, 2026 | Automatic Metrics | Pairwise Preference | Accuracy | Not reported |
| TwinKV: A Composable Repair Pass for KV Cache Eviction via Pairwise Key Redundancy Aug 27, 2026 | Automatic Metrics | Pairwise Preference | Spearman | Not reported |
| Memory Augmentation Unlocks Efficient Chain-of-Thought Reasoning Aug 21, 2026 | Automatic Metrics | Demonstrations | Accuracy, Latency | Not reported |
| Will Scaling Improve Social Simulation with LLMs? Jul 2, 2026 | Automatic Metrics, Simulation Env | Not reported | Accuracy | Calibration |
| Cliff Tokens: Identifying Single-Token Failure Triggers in LLM Mathematical Reasoning Jun 24, 2026 | Automatic Metrics | Pairwise Preference | Accuracy, Pass@64 | Not reported |
| Enhanced LLM Reasoning by Optimizing Reward Functions with Search-Driven Reinforcement Learning May 3, 2026 | Automatic Metrics | Pairwise Preference | Accuracy, F1 | Not reported |
| Hidden Measurement Error in LLM Pipelines Distorts Annotation, Evaluation, and Benchmarking Apr 13, 2026 | Llm As Judge | Demonstrations | Precision, Agreement | Not reported |
| Auditing MCQA Benchmarks through Probability Landscapes Aug 31, 2026 | Not reported | Pairwise Preference | Not reported | Not reported |
| RARE: Decoupling Representation Steering from Expert Routing in Mixture-of-Experts Language Models Aug 21, 2026 | Automatic Metrics | Not reported | Accuracy, Success rate | Not reported |
| Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization Aug 20, 2026 | Automatic Metrics | Not reported | Accuracy | Not reported |
Gap: Human feedback
Human feedback is present in 8 of 60 papers.
Gap: Quality controls
Quality controls is present in 2 of 60 papers.
Strong: Benchmarks
Benchmarks is present in 60 of 60 papers.
Strong: Metrics
Metrics is present in 40 of 60 papers.
Gap: Known rater population
Known rater population is present in 6 of 60 papers.
Gap: Known annotation unit
Known annotation unit is present in 10 of 60 papers.
Evaluation Modes
Human Feedback Mix
Top Benchmarks
Top Metrics
Lizhuo Zhang, Mengmeng Tang, Chenfeng Long, Xiaoyong Tang, Xiang Luo · Aug 19, 2026 · Citations: 0
The mechanical reference is leak-free: each case's preference and accuracy are estimated from its other runs only.
Minsoo Song, Chanjun Park · Aug 31, 2026 · Citations: 0
To provide a scalable diagnostic approach, we propose a two-component probabilistic framework for auditing MCQA benchmarks using model output distributions.
Dhruv Deshmukh, Saurabh Goyal, Nipun Kwatra, Ramachandran Ramjee · Dec 18, 2025 · Citations: 0
Kascade achieves up to 4.1x speedup in decode attention and 2.2x speedup in prefill attention over FlashAttention-3 baseline on H100 GPUs while closely matching dense attention accuracy on long-context benchmarks such as LongBench and…
Hong Chen, Yudong Zeng, Yongwei Huang, Zuhao Ouyang, Dongnan Zheng · Aug 27, 2026 · Citations: 0
We introduce TwinKV, a training-free, attention-free redundancy signal that detects whether a token's key has a near-duplicate elsewhere in context.
Ruotong Liao, Nikolai Röhrich, Xiaohan Wang, Yuhui Zhang, Yasaman Samadzadeh · Mar 2, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Haodong Zhao, Jidong Li, Zhaomin Wu, Tianjie Ju, Zhuosheng Zhang · Sep 25, 2025 · Citations: 0
Understanding persuasion is critical for the safety and reliability of multi-agent systems built on large language models (LLMs).
Simeng Zhang, Yilong Chen, Wenyuan Zhang, Zhenyu Zhang, Yao Chen · Aug 21, 2026 · Citations: 0
Based on this principle, we propose Memory-Augmented Compression, a training-free framework that constructs reusable reasoning memories from historical traces and retrieves them as prefill-side scaffolds.
Zhibo Zhang, Zhen Ouyang, Ling Shi, Kailong Wang · Aug 21, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Qian Kou, Xiaofeng Shi, Xiaosong Qiu, Hua Zhou · Aug 20, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Haonan He, Xinyue Fan · Aug 20, 2026 · Citations: 0
For instance, LoRA-GA^2 surpasses the leading baseline by an average of 0.66 points on the GLUE benchmark, and outperforms the strongest baseline by 1.03 points on GSM8K and 0.87 points on HumanEval, respectively.
Sarah Breckner, Sebastian Schuster · Mar 5, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Zhixuan Liu, Zhichen Dong, Yuanfu Wang, Chao Yang · Aug 13, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Atahan Dokme, Benjamin Reichman, Larry Heck · Apr 9, 2026 · Citations: 0
Beyond emotion specifically, the benchmark construction procedure offers a controllable instrument for stylistic translation and robustness evaluation.
Shailja Thakur, Sungeun An, Chad DeLuca, Hima Patel · Aug 12, 2026 · Citations: 0
A benchmark score comes from a single phrasing of each problem.
Zhenyu Zhao, Sander Land, Daniel M. Bikel, Waseem Alshikh · Apr 29, 2026 · Citations: 0
Across three model families and five mathematical reasoning benchmarks, our approach shortens reasoning traces by 8.1% on average; under a TOST equivalence analysis at a +/- 2pp margin, accuracy is equivalent or inconclusive on 13/15 model…
Caleb Ziems, William Held, Su Doga Karaca, David Grusky, Tatsunori Hashimoto · Jul 2, 2026 · Citations: 0
We use scaling laws to study the relationship between LLMs' compute scale, general capability benchmarks, and the fidelity of social simulation in three representative sub-domains: opinion modeling, behavioral simulation, and longitudinal…
Xiangdong Zhang, Debing Zhang, Shaofeng Zhang, Xiaohan Qin, Yu Cheng · May 24, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Andikawati P Widjaja, Yongjun Kim, Hyounghun Kim, Jaeho Lee · Jul 2, 2026 · Citations: 0
Across eight benchmarks (including MMLU, GSM8K, and RULER) and three model families (Qwen2.5, Llama3.2, Gemma4), PartRep retains most of the gains of full repetition while using only 59.4\% of its KV cache and 79.0\% of its prefill FLOPs.
Kevin Yandoka Denamganaï · May 27, 2026 · Citations: 0
Self-evolving scientific agents capable of conquering the hard tail of formal mathematics require Compositional Learning Behaviours (CLBs) -- the capacity to ground and recombine novel symbolic structures in context, beyond mere…
Subhadip Mitra · Jun 28, 2026 · Citations: 0
Inference-time safety methods for large language models have proliferated, yet no systematic comparison exists.
Steven Kolawole, Virginia Smith · Jun 25, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Avinash Reddy, Thayne T. Walker, James S. Ide, Amrit Singh Bedi · Feb 8, 2026 · Citations: 0
Across structured reasoning benchmarks, DCCD improves strict structured accuracy by up to +24 percentage points over standard constrained decoding (e.g., 15.2\% to 39.0\% on GSM8K with a 1B model), and enables smaller model pairs to match…
Xiangyue Liu, Zijian Zhang, Miles Yang, Zhao Zhong, Liefeng Bo · Apr 9, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Hung-Hsuan Chen · Feb 26, 2026 · Citations: 0
On the SlimOrca benchmark, CeRA breaks this linear barrier: at rank 64 (PPL 3.89), it outperforms LoRA at rank 512 (PPL 3.90), demonstrating superior spectral efficiency.
Jaeyong Ko, Pilsung Kang, Yukyung Lee · Jun 24, 2026 · Citations: 0
Across seven models and three mathematical reasoning benchmarks (GSM1K, MATH500, AIME 2025), cliff tokens act as failure triggers; deleting the first cliff token and resampling recovers pass@64 to 1.0, while keeping it limits recovery to…
Tianyu Dong, Yangyang Liu, Jiang Zhou, Xinwei Wu, Xiaohu Zhao · Jun 24, 2026 · Citations: 0
We conduct experiments on 2 LLMs across 5 low-resource languages and 3 benchmarks.
Azher Ali, Ibtsam Haider, Raja Khurram Shahzad, Seemab Latif, Mehwish Fatima · Jun 24, 2026 · Citations: 0
Recent LLMs demonstrate strong mathematical reasoning capabilities, but existing gains rely heavily on English-centric training resources and benchmarks.
Chenhao Dang, Jing Ma, Mingjie Liao · Jun 23, 2026 · Citations: 0
On The Pile benchmark, HDS reaches the final validation perplexity of the next best method with 44% fewer training iterations.
Fengfeng Liang, Yuechen Zhang, Jiaya Jia · Jun 23, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Yu Deng · Jun 18, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Ali Asgarov, Umid Suleymanov, Aadyant Khatri · Oct 31, 2025 · Citations: 0
We introduce SIGMA (Search-Augmented On-Demand Knowledge Integration for AGentic Mathematical reAsoning), a unified framework that orchestrates specialized agents to independently reason, perform targeted searches, and synthesize findings…
Glenn Matlin, Chandreyi Chakraborty, Saehee Eom, Mika Okamoto, Rayan Castilla · Jun 17, 2026 · Citations: 0
Training-data attribution measures how strongly each training document influences a model's predictions on a benchmark, but document-level scores are too noisy to identify which corpus regions support which capabilities, and prior work has…
Anany Kotawala · May 28, 2026 · Citations: 0
Across two public LLM leaderboards, many displayed pairwise rankings do not meet a conventional paired-test resolution target under the actual paired evaluation design: 11 of 40 Open LLM Leaderboard v1 pairwise comparisons and 4 of 9…
Siddharth Boppana, Annabel Ma, Max Loeffler, Raphael Sarfati, Eric Bigelow · Mar 5, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Tanmoy Chakraborty, Ayan Sengupta, Suparna Bhattacharya, Partha Pratim Chakrabarti, Amlan Chakrabarti · May 28, 2026 · Citations: 0
Large language models (LLMs) frequently achieve impressive scores on standardized benchmarks, yet accuracy alone offers a limited view of their capabilities.
Dominika Agnieszka Długosz, Arlindo Oliveira, Natalia Díaz-Rodríguez · May 27, 2026 · Citations: 0
The GSM-Symbolic benchmark (Mirzadeh et al., 2025) reported consistent performance drops across 25 Large Language Models (LLMs) when tested on template-generated variants of GSM8K problems, concluding that the models lack genuine reasoning…
G M Shahariar, Erfan Shayegani, Ali Nazari, Nael Abu-Ghazaleh · Oct 25, 2025 · Citations: 0
Experiments across four benchmarks (AIME25, MATH-500, GSM8K, and GPQA Diamond) using three state-of-the-art open reasoning models demonstrate that Q-Value steering policy achieves significant performance gains with "surgical" efficiency,…
Woomin Song, Saket Dingliwal, Sai Muralidhar Jayanthi, Bhavana Ganesh, Jinwoo Shin · Jun 5, 2025 · Citations: 0
Extensive evaluations across multiple models and reasoning tasks (AIME-2024, GPQA-Diamond, and LiveCodeBench) demonstrate that STAND reduces inference latency by 60-65% compared to standard autoregressive decoding while maintaining…
Sayed Mohammadreza Tayaranian Hosseini, Amir Ardakani, Warren J. Gross · Feb 26, 2026 · Citations: 0
Our evaluation experiments on Llama models shows that InnerQ maintains a few-shot GSM8K performance comparable to non-quantized KV caches and surpasses prior KV cache quantization methods.
Andrea Sassella, Andrea Chizzola, Tommaso Bianchi, Luca Alessandrelli, Mark James Carman · May 8, 2026 · Citations: 0
This report benchmarks the performance of ENGINEERING Ingegneria Informatica S.p.A.'s EngGPT2MoE-16B-A3B LLM, a 16B parameter Mixture of Experts (MoE) model with 3B active parameters.
Arash Ahmadi, Sarah Sharif, Yaser, Banad · May 3, 2026 · Citations: 0
Mathematical reasoning is a key benchmark for large language models.
Nils Grünefeld, Bertram Højer, Philipp Mondorf, Barbara Plank, Anna Rogers · May 8, 2026 · Citations: 0
Language model (LM) "reasoning", commonly described as Chain-of-Thought or test-time scaling, often improves benchmark performance, but the dynamics underlying this process remain poorly understood.
Qiyong Zhong, Mao Zheng, Mingyang Song, Xin Lin, Jie Sun · May 8, 2026 · Citations: 0
To address this, we propose SOD, a step-wise on-policy distillation framework for small language model agents, which adaptively reweights distillation strength at each step based on step-level divergence.
Tsuyoshi Okita · May 8, 2026 · Citations: 0
With a 7B-parameter LLM whose weights are entirely frozen, CIKA achieves 69.7\% on the contamination-free Omni-MATH-Rule benchmark and 64.0\% overall, compared to 60.5\% for o1-mini, and 97.2\% on GSM8K, 46--50\% on AIME 2024--2026, and…
Qianjia Cheng, Yuchen Zhang, Zhilin Wang, Yuxin Zuo, Shunkai Zhang · May 7, 2026 · Citations: 0
Paradoxically, we observe that tool-enabled evaluation can degrade reasoning performance even when the strong thinking models make almost no actual tool calls.
Solomon Messing · Apr 13, 2026 · Citations: 0
LLM evaluations drive which models get deployed, what safety standards get adopted, which research conclusions get published, and how projections of AI's labor-market impact get made.
Xiaoyu Xu, Minxin Du, Kun Fang, Yaxin Xiao, Zhicong Huang · Jan 29, 2026 · Citations: 0
Furthermore, to facilitate rigorous evaluation, we introduce PCH, a unified benchmark encompassing Personal, Copyrighted, and Harmful content, alongside two symmetric metrics, Forget Degree (F.D.) and Retain Utility (R.U.), to…
Buu Phan, Ashish Khisti, Karen Ullrich · Dec 16, 2025 · Citations: 0
Our method enables sequence likelihood evaluation for vocabularies different from the teacher model native tokenizer, addressing two specific scenarios: when the student vocabulary is a subset of the teacher vocabulary, and the general case…
Weiqin Wang, Yile Wang, Kehao Chen, Hui Huang · Dec 17, 2025 · Citations: 0
We conduct experiments across various models and benchmarks, experimental results show that SCOPE consistently outperforms recent baselines.
Pere Martra · Dec 27, 2025 · Citations: 0
We evaluated seven expansion ratio configurations using comprehensive benchmarks assessing factual knowledge, mathematical reasoning, language comprehension, instruction-following, and truthfulness.
Weitao Li, Boran Xiang, Xiaolong Wang, Zhinan Gou, Weizhi Ma · Aug 8, 2025 · Citations: 0
Experiments on open-domain QA, MMLU-Pro, medical, and mathematical reasoning tasks show that UR^2, built on Qwen-2.5-3/7B and LLaMA-3.1-8B, consistently outperforms existing RAG and RL baselines, and achieves performance comparable to…
Yutao Hou, Zeguan Xiao, Fei Yu, Yihan Jiang, Ma Shuguang · Jun 5, 2025 · Citations: 0
Existing robustness evaluations predominantly rely on hand-crafted templates or a limited set of perturbation rules.
Aofan Liu, Jingxiang Meng · Apr 24, 2026 · Citations: 0
Iterative self-correction is widely used in agentic LLM systems, but when repeated refinement helps versus hurts remains unclear.
Yijia Fan, Jusheng Zhang, Kaitong Cai, Jing Yang, Chengpei Tang · Nov 17, 2025 · Citations: 0
To address this, we introduce the Dynamic Auction-based Language Agent (DALA), a novel framework that treats communication bandwidth as a scarce and tradable resource.
Hector Borobia, Elies Seguí-Mas, Guillermina Tormo-Carbó · Apr 24, 2026 · Citations: 0
We systematically study component-type LoRA placement across two hybrid architectures -- Qwen3.5-0.8B (sequential, GatedDeltaNet + softmax attention) and Falcon-H1-0.5B (parallel, Mamba-2 SSM + attention) -- fine-tuned on three domains and…
John Kirchenbauer, Abhimanyu Hans, Brian Bartoldson, Micah Goldblum, Ashwinee Panda · Feb 5, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Shangqing Tu, Yaxuan Li, Yushi Bai, Lei Hou, Juanzi Li · Oct 9, 2025 · Citations: 0
Our method features a specialized judge model trained with out-of-distribution data (AIME 2022, AIME 2023, and MATH 500) using oversampling techniques to accurately predict answer equivalence from partial reasoning traces, achieving 0.7072…
Teng Wang, Zhangyi Jiang, Zhenqi He, Shenyang Tong, Wenhan Yang · Mar 16, 2025 · Citations: 0
Empirical results on the PRM800K dataset show that HRM, together with HNC, provides more stable and reliable evaluations than PRM.
Khushal Sethi · Apr 9, 2026 · Citations: 0
We introduce TrACE (Trajectorical Adaptive Compute via agrEement), a training-free controller that allocates LLM calls adaptively across agent timesteps by measuring inter-rollout action agreement.
Pengxiang Li, Yefan Zhou, Dilxat Muhtar, Lu Yin, Shilin Yan · Aug 27, 2025 · Citations: 0
Empirical evaluations of LLaDA-8B and Dream-7B across multiple tasks show that Prophet reduces the number of decoding steps by up to 3.4x while preserving high generation quality.