HFEPX Benchmark Hub
AIME or HotpotQA or MMLU Benchmark Papers
Updated from current HFEPX corpus (2026-09-07). This page tracks 60 papers reporting AIME or HotpotQA or MMLU benchmark evidence, with protocol and metric context for comparison.
HFEPX Benchmark Hub
Updated from current HFEPX corpus (2026-09-07). This page tracks 60 papers reporting AIME or HotpotQA or MMLU benchmark evidence, with protocol and metric context for comparison.
Use this page for benchmark-matched method comparisons and eval protocol selection. Quality band: High .
High-Signal Coverage
100.0%
60 / 60 sampled papers are not low-signal flagged.
Replication-Ready Set
41
Papers with explicit benchmark + metric + eval mode fields.
Quality Controls
3.3%
2 papers report calibration/adjudication/IAA controls.
Primary action: Start with the top 2 benchmark-matched papers, then compare evaluation modes in the protocol matrix.
Ranked by protocol completeness so you can quickly find papers suitable for comparison studies.
Aug 19, 2026 · Citations: 0 · Score: 8.5
Eval: Automatic Metrics · Metrics: Accuracy
Aug 27, 2026 · Citations: 0 · Score: 8.5
Eval: Automatic Metrics · Metrics: Spearman
Aug 21, 2026 · Citations: 0 · Score: 8.5
Eval: Automatic Metrics · Metrics: Accuracy
Jul 2, 2026 · Citations: 0 · Score: 8.0
Eval: Automatic Metrics, Simulation Env · Metrics: Accuracy
Jun 24, 2026 · Citations: 0 · Score: 8.0
Eval: Automatic Metrics · Metrics: Accuracy
May 6, 2026 · Citations: 0 · Score: 7.5
Eval: Automatic Metrics · Metrics: F1
Compare protocol ingredients quickly before deep-reading full papers.
| Paper | Eval Modes | Human Feedback | Metrics | Quality Controls |
|---|---|---|---|---|
| Decomposing Wrong-Consensus Agreement in LLM Self-Consistency Aug 19, 2026 | Automatic Metrics | Pairwise Preference | Accuracy | Not reported |
| TwinKV: A Composable Repair Pass for KV Cache Eviction via Pairwise Key Redundancy Aug 27, 2026 | Automatic Metrics | Pairwise Preference | Spearman | Not reported |
| Memory Augmentation Unlocks Efficient Chain-of-Thought Reasoning Aug 21, 2026 | Automatic Metrics | Demonstrations | Accuracy, Latency | Not reported |
| Will Scaling Improve Social Simulation with LLMs? Jul 2, 2026 | Automatic Metrics, Simulation Env | Not reported | Accuracy | Calibration |
| Cliff Tokens: Identifying Single-Token Failure Triggers in LLM Mathematical Reasoning Jun 24, 2026 | Automatic Metrics | Pairwise Preference | Accuracy, Pass@64 | Not reported |
| CAR: Query-Guided Confidence-Aware Reranking for Retrieval-Augmented Generation May 6, 2026 | Automatic Metrics | Pairwise Preference | F1, Ndcg | Not reported |
| Hidden Measurement Error in LLM Pipelines Distorts Annotation, Evaluation, and Benchmarking Apr 13, 2026 | Llm As Judge | Demonstrations | Precision, Agreement | Not reported |
| RARE: Decoupling Representation Steering from Expert Routing in Mixture-of-Experts Language Models Aug 21, 2026 | Automatic Metrics | Not reported | Accuracy, Success rate | Not reported |
| Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization Aug 20, 2026 | Automatic Metrics | Not reported | Accuracy | Not reported |
| From Retrieved Context to Runtime Control: Adaptive Compression for Edge-based RAG Aug 20, 2026 | Not reported | Not reported | Latency | Not reported |
Gap: Human feedback
Human feedback is present in 9 of 60 papers.
Gap: Quality controls
Quality controls is present in 2 of 60 papers.
Strong: Benchmarks
Benchmarks is present in 60 of 60 papers.
Strong: Metrics
Metrics is present in 44 of 60 papers.
Gap: Known rater population
Known rater population is present in 6 of 60 papers.
Gap: Known annotation unit
Known annotation unit is present in 12 of 60 papers.
Evaluation Modes
Human Feedback Mix
Top Benchmarks
Top Metrics
Lizhuo Zhang, Mengmeng Tang, Chenfeng Long, Xiaoyong Tang, Xiang Luo · Aug 19, 2026 · Citations: 0
The mechanical reference is leak-free: each case's preference and accuracy are estimated from its other runs only.
Minsoo Song, Chanjun Park · Aug 31, 2026 · Citations: 0
To provide a scalable diagnostic approach, we propose a two-component probabilistic framework for auditing MCQA benchmarks using model output distributions.
Yang Zhao, Chengxiao Dai, Mengying Kou, Yue Xiu, Dusit Niyato · Jan 29, 2026 · Citations: 0
We present SHARD-MEMO, an agentic memory system built on scope-before-routing: metadata predicates first identify the admissible shards, and a learned router then selects a small number of them for shard-local approximate nearest neighbor…
Dhruv Deshmukh, Saurabh Goyal, Nipun Kwatra, Ramachandran Ramjee · Dec 18, 2025 · Citations: 0
Kascade achieves up to 4.1x speedup in decode attention and 2.2x speedup in prefill attention over FlashAttention-3 baseline on H100 GPUs while closely matching dense attention accuracy on long-context benchmarks such as LongBench and…
Hong Chen, Yudong Zeng, Yongwei Huang, Zuhao Ouyang, Dongnan Zheng · Aug 27, 2026 · Citations: 0
We introduce TwinKV, a training-free, attention-free redundancy signal that detects whether a token's key has a near-duplicate elsewhere in context.
Ruotong Liao, Nikolai Röhrich, Xiaohan Wang, Yuhui Zhang, Yasaman Samadzadeh · Mar 2, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Haodong Zhao, Jidong Li, Zhaomin Wu, Tianjie Ju, Zhuosheng Zhang · Sep 25, 2025 · Citations: 0
Understanding persuasion is critical for the safety and reliability of multi-agent systems built on large language models (LLMs).
Simeng Zhang, Yilong Chen, Wenyuan Zhang, Zhenyu Zhang, Yao Chen · Aug 21, 2026 · Citations: 0
Based on this principle, we propose Memory-Augmented Compression, a training-free framework that constructs reusable reasoning memories from historical traces and retrieves them as prefill-side scaffolds.
Zhibo Zhang, Zhen Ouyang, Ling Shi, Kailong Wang · Aug 21, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Qian Kou, Xiaofeng Shi, Xiaosong Qiu, Hua Zhou · Aug 20, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Zlatan Feric, Amir Taherin, Yanzhi Wang, David Kaeli · Aug 20, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Weimeng Luo · Aug 13, 2026 · Citations: 0
We adapt S2G-RAG's structured sufficiency-and-gap judgment to a frozen Search-R1 pipeline and train a Qwen3.5-2B judge on 3,009 states from 900 disjoint HotpotQA questions.
Xinlong Xu, Yoshua Y. Li · Aug 13, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Zhixuan Liu, Zhichen Dong, Yuanfu Wang, Chao Yang · Aug 13, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Zhipeng Song, Yizhi Zhou, Xiangyu Kong, Jiulong Jiao, Xuezhou Ye · May 6, 2026 · Citations: 0
CAR converts these confidence changes into coarse precedence constraints and returns the feasible ranking with minimum Kendall distance from the baseline, preserving existing pairwise preferences unless generator-side evidence supports…
Yuchao Wu, Junqin Li, XingCheng Liang, Yongjie Chen, Yinghao Liang · Aug 12, 2026 · Citations: 0
Experiments on HotpotQA, 2WikiMultiHopQA, and MuSiQue show that SAG achieves the best retrieval and end-to-end QA performance on every benchmark, with gains that widen as reasoning-chain complexity increases.
Tal Oved, Roi Pony, Oshri Naparstek, Udi barzelay · Aug 11, 2026 · Citations: 0
Evolutionary optimization of LLM prompts and agentic programs (e.g., GEPA) is dominated by fitness evaluation: scoring each candidate runs an answering LLM over a validation set, so the evaluator's price tier dictates total search cost.
Shailja Thakur, Sungeun An, Chad DeLuca, Hima Patel · Aug 12, 2026 · Citations: 0
A benchmark score comes from a single phrasing of each problem.
Zhenyu Zhao, Sander Land, Daniel M. Bikel, Waseem Alshikh · Apr 29, 2026 · Citations: 0
Across three model families and five mathematical reasoning benchmarks, our approach shortens reasoning traces by 8.1% on average; under a TOST equivalence analysis at a +/- 2pp margin, accuracy is equivalent or inconclusive on 13/15 model…
Caleb Ziems, William Held, Su Doga Karaca, David Grusky, Tatsunori Hashimoto · Jul 2, 2026 · Citations: 0
We use scaling laws to study the relationship between LLMs' compute scale, general capability benchmarks, and the fidelity of social simulation in three representative sub-domains: opinion modeling, behavioral simulation, and longitudinal…
Xiangdong Zhang, Debing Zhang, Shaofeng Zhang, Xiaohan Qin, Yu Cheng · May 24, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Andikawati P Widjaja, Yongjun Kim, Hyounghun Kim, Jaeho Lee · Jul 2, 2026 · Citations: 0
Across eight benchmarks (including MMLU, GSM8K, and RULER) and three model families (Qwen2.5, Llama3.2, Gemma4), PartRep retains most of the gains of full repetition while using only 59.4\% of its KV cache and 79.0\% of its prefill FLOPs.
Ananto Nayan Bala · Jul 1, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Kevin Yandoka Denamganaï · May 27, 2026 · Citations: 0
Self-evolving scientific agents capable of conquering the hard tail of formal mathematics require Compositional Learning Behaviours (CLBs) -- the capacity to ground and recombine novel symbolic structures in context, beyond mere…
Subhadip Mitra · Jun 28, 2026 · Citations: 0
Inference-time safety methods for large language models have proliferated, yet no systematic comparison exists.
Steven Kolawole, Virginia Smith · Jun 25, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Xiangyue Liu, Zijian Zhang, Miles Yang, Zhao Zhong, Liefeng Bo · Apr 9, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Binqi Shen, Lier Jin, Hanyu Cai, Lan Hu, Yuting Xin · May 21, 2026 · Citations: 0
Unlike existing evaluations that compare methods in isolation, the proposed framework enables decision-oriented analysis.
Jaeyong Ko, Pilsung Kang, Yukyung Lee · Jun 24, 2026 · Citations: 0
Across seven models and three mathematical reasoning benchmarks (GSM1K, MATH500, AIME 2025), cliff tokens act as failure triggers; deleting the first cliff token and resampling recovers pass@64 to 1.0, while keeping it limits recovery to…
Tianyu Dong, Yangyang Liu, Jiang Zhou, Xinwei Wu, Xiaohu Zhao · Jun 24, 2026 · Citations: 0
We conduct experiments on 2 LLMs across 5 low-resource languages and 3 benchmarks.
Chenhao Dang, Jing Ma, Mingjie Liao · Jun 23, 2026 · Citations: 0
On The Pile benchmark, HDS reaches the final validation perplexity of the next best method with 44% fewer training iterations.
Fengfeng Liang, Yuechen Zhang, Jiaya Jia · Jun 23, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Ali Asgarov, Umid Suleymanov, Aadyant Khatri · Oct 31, 2025 · Citations: 0
We introduce SIGMA (Search-Augmented On-Demand Knowledge Integration for AGentic Mathematical reAsoning), a unified framework that orchestrates specialized agents to independently reason, perform targeted searches, and synthesize findings…
Glenn Matlin, Chandreyi Chakraborty, Saehee Eom, Mika Okamoto, Rayan Castilla · Jun 17, 2026 · Citations: 0
Training-data attribution measures how strongly each training document influences a model's predictions on a benchmark, but document-level scores are too noisy to identify which corpus regions support which capabilities, and prior work has…
Anany Kotawala · May 28, 2026 · Citations: 0
Across two public LLM leaderboards, many displayed pairwise rankings do not meet a conventional paired-test resolution target under the actual paired evaluation design: 11 of 40 Open LLM Leaderboard v1 pairwise comparisons and 4 of 9…
Siddharth Boppana, Annabel Ma, Max Loeffler, Raphael Sarfati, Eric Bigelow · Mar 5, 2026 · Citations: 0
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
Tanmoy Chakraborty, Ayan Sengupta, Suparna Bhattacharya, Partha Pratim Chakrabarti, Amlan Chakrabarti · May 28, 2026 · Citations: 0
Large language models (LLMs) frequently achieve impressive scores on standardized benchmarks, yet accuracy alone offers a limited view of their capabilities.
Woomin Song, Saket Dingliwal, Sai Muralidhar Jayanthi, Bhavana Ganesh, Jinwoo Shin · Jun 5, 2025 · Citations: 0
Extensive evaluations across multiple models and reasoning tasks (AIME-2024, GPQA-Diamond, and LiveCodeBench) demonstrate that STAND reduces inference latency by 60-65% compared to standard autoregressive decoding while maintaining…
Andrea Sassella, Andrea Chizzola, Tommaso Bianchi, Luca Alessandrelli, Mark James Carman · May 8, 2026 · Citations: 0
This report benchmarks the performance of ENGINEERING Ingegneria Informatica S.p.A.'s EngGPT2MoE-16B-A3B LLM, a 16B parameter Mixture of Experts (MoE) model with 3B active parameters.
Tianyi Huang, Samuel Xu, Jason Tansong Dang, Samuel Yan, Kimberley Yin · Apr 19, 2026 · Citations: 0
Agentic systems often fail not by being entirely wrong, but by being too precise: a response may be generally useful while particular claims exceed what the evidence supports.
Qiyong Zhong, Mao Zheng, Mingyang Song, Xin Lin, Jie Sun · May 8, 2026 · Citations: 0
To address this, we propose SOD, a step-wise on-policy distillation framework for small language model agents, which adaptively reweights distillation strength at each step based on step-level divergence.
Tsuyoshi Okita · May 8, 2026 · Citations: 0
With a 7B-parameter LLM whose weights are entirely frozen, CIKA achieves 69.7\% on the contamination-free Omni-MATH-Rule benchmark and 64.0\% overall, compared to 60.5\% for o1-mini, and 97.2\% on GSM8K, 46--50\% on AIME 2024--2026, and…
Zhichao Yan, Yunxiao Zhao, Jiapu Wang, Jiaoyan Chen, Xiaoli Li · Jan 21, 2026 · Citations: 0
Current evaluation methods for Retrieval Augmented Generation (RAG) suffer from factual myopia: they relentlessly emphasize factual accuracy yet neglect global logical integrity in long-form answer generation.
Qianjia Cheng, Yuchen Zhang, Zhilin Wang, Yuxin Zuo, Shunkai Zhang · May 7, 2026 · Citations: 0
Paradoxically, we observe that tool-enabled evaluation can degrade reasoning performance even when the strong thinking models make almost no actual tool calls.
Solomon Messing · Apr 13, 2026 · Citations: 0
LLM evaluations drive which models get deployed, what safety standards get adopted, which research conclusions get published, and how projections of AI's labor-market impact get made.
Xiaoyu Xu, Minxin Du, Kun Fang, Yaxin Xiao, Zhicong Huang · Jan 29, 2026 · Citations: 0
Furthermore, to facilitate rigorous evaluation, we introduce PCH, a unified benchmark encompassing Personal, Copyrighted, and Harmful content, alongside two symmetric metrics, Forget Degree (F.D.) and Retain Utility (R.U.), to…
Weiqin Wang, Yile Wang, Kehao Chen, Hui Huang · Dec 17, 2025 · Citations: 0
We conduct experiments across various models and benchmarks, experimental results show that SCOPE consistently outperforms recent baselines.
Pere Martra · Dec 27, 2025 · Citations: 0
We evaluated seven expansion ratio configurations using comprehensive benchmarks assessing factual knowledge, mathematical reasoning, language comprehension, instruction-following, and truthfulness.
Matthew Penaroza · Apr 8, 2026 · Citations: 0
Language model agents reason from scratch on every query, discarding their chain of thought after each run.
Jingxi Qiu, Zeyu Han, Cheng Huang · May 5, 2026 · Citations: 0
We evaluate on HotpotQA-RAG v3, a controlled multi-hop benchmark, under an artifact-aware protocol (shortcut baselines, counterfactual swaps, no-oracle checks, GPT-4o audits).
Weitao Li, Boran Xiang, Xiaolong Wang, Zhinan Gou, Weizhi Ma · Aug 8, 2025 · Citations: 0
Experiments on open-domain QA, MMLU-Pro, medical, and mathematical reasoning tasks show that UR^2, built on Qwen-2.5-3/7B and LLaMA-3.1-8B, consistently outperforms existing RAG and RL baselines, and achieves performance comparable to…
Yijia Fan, Jusheng Zhang, Kaitong Cai, Jing Yang, Chengpei Tang · Nov 17, 2025 · Citations: 0
To address this, we introduce the Dynamic Auction-based Language Agent (DALA), a novel framework that treats communication bandwidth as a scarce and tradable resource.
Ruiyi Yang, Hao Xue, Imran Razzak, Hakim Hacid, Flora D. Salim · Oct 23, 2025 · Citations: 0
A Head Agent provides guidance that leads retrieval, while an Iteration Agent selects and expands HSeq via structure-respecting actions (e.g., parent/child hops, table row/column neighbors, KG relations); Finally the head agent composes…
Ziliang Wang, Kang An, Xuhui Zheng, Faqiang Qian, Weikun Zhang · Oct 1, 2025 · Citations: 0
We propose Erasable Reinforcement Learning (ERL), a novel framework that transforms fragile reasoning into a robust process.
Shangqing Tu, Yaxuan Li, Yushi Bai, Lei Hou, Juanzi Li · Oct 9, 2025 · Citations: 0
Our method features a specialized judge model trained with out-of-distribution data (AIME 2022, AIME 2023, and MATH 500) using oversampling techniques to accurately predict answer equivalence from partial reasoning traces, achieving 0.7072…
Pengxiang Li, Yefan Zhou, Dilxat Muhtar, Lu Yin, Shilin Yan · Aug 27, 2025 · Citations: 0
Empirical evaluations of LLaDA-8B and Dream-7B across multiple tasks show that Prophet reduces the number of decoding steps by up to 3.4x while preserving high generation quality.
Zhichao Wang · Oct 27, 2025 · Citations: 0
This paper proposes Group-relative Implicit Fine-Tuning (GIFT), a reinforcement learning framework for aligning large language models (LLMs) that unifies on-policy optimization with implicit preference learning.
Ziyi Liu · Sep 22, 2025 · Citations: 0
Our work offers an effective solution for optimizing LLMs in long-range interactions, providing new insights for developing more robust Agents.
Shu Wang, Edwin Yu, Oscar Love, Tom Zhang, Tom Wong · Apr 6, 2026 · Citations: 0
Large Language Model (LLM) agents require persistent memory to maintain personalization, factual continuity, and long-horizon reasoning, yet standard context-window and retrieval-augmented generation (RAG) pipelines degrade over…
Yizhou Liu, Qi Sun, Yulin Chen, Siyue Zhang, Chen Zhao · Apr 6, 2026 · Citations: 0
Agents equipped with search tools have emerged as effective solutions for knowledge-intensive tasks.