Skip to content
OpenTrain AIFor AI Companies
← Back to explorer

HFEPX Benchmark Hub

AIME or AlpacaEval or MMLU Benchmark Papers

Updated from current HFEPX corpus (2026-08-30). This page tracks 60 papers reporting AIME or AlpacaEval or MMLU benchmark evidence, with protocol and metric context for comparison.

Papers: 60 Last published: Aug 21, 2026 Global RSS

Researcher Quick Triage

Use this page for benchmark-matched method comparisons and eval protocol selection. Quality band: High .

High-Signal Coverage

100.0%

60 / 60 sampled papers are not low-signal flagged.

Replication-Ready Set

34

Papers with explicit benchmark + metric + eval mode fields.

Quality Controls

5.0%

3 papers report calibration/adjudication/IAA controls.

  • 60 papers explicitly name benchmark datasets in the sampled set.
  • 37 papers report at least one metric term in metadata extraction.
  • Start with the ranked shortlist below before reading all papers.

Primary action: Start with the top 2 benchmark-matched papers, then compare evaluation modes in the protocol matrix.

Why This Matters (Expanded)

Why This Matters For Eval Research

  • Use this page to compare AIME or AlpacaEval or MMLU papers by evaluation mode, metric, and evidence quality before reusing reported results.
Protocol Notes (Expanded)

Protocol Takeaways

  • AIME or AlpacaEval or MMLU papers are often paired with automatic_metrics, llm_as_judge.

Benchmark Interpretation

  • MMLU: 39 papers
  • AIME: 17 papers
  • GSM8K: 12 papers
  • MATH-500: 6 papers

Metric Interpretation

  • accuracy: 20 papers
  • cost: 8 papers
  • latency: 4 papers
  • perplexity: 3 papers

Start Here (Benchmark-Matched First 6)

Ranked by protocol completeness so you can quickly find papers suitable for comparison studies.

Protocol Matrix (Top 10)

Compare protocol ingredients quickly before deep-reading full papers.

Paper Eval Modes Human Feedback Metrics Quality Controls
Memory Augmentation Unlocks Efficient Chain-of-Thought Reasoning

Aug 21, 2026

Automatic Metrics Demonstrations Accuracy, Latency Not reported
Will Scaling Improve Social Simulation with LLMs?

Jul 2, 2026

Automatic Metrics, Simulation Env Not reported Accuracy Calibration
Cliff Tokens: Identifying Single-Token Failure Triggers in LLM Mathematical Reasoning

Jun 24, 2026

Automatic Metrics Pairwise Preference Accuracy, Pass@64 Not reported
Hidden Measurement Error in LLM Pipelines Distorts Annotation, Evaluation, and Benchmarking

Apr 13, 2026

Llm As Judge Demonstrations Precision, Agreement Not reported
Diagnosing Translated Benchmarks: An Automated Quality Assurance Study of the EU20 Benchmark Suite

Apr 2, 2026

Automatic Metrics Not reported Accuracy Calibration, Gold Questions
PubMed Reasoner: Dynamic Reasoning-based Retrieval for Evidence-Grounded Biomedical Question Answering

Mar 28, 2026

Llm As Judge, Automatic Metrics Expert Verification Accuracy, Relevance Not reported
DSPA: Dynamic SAE Steering for Data-Efficient Preference Alignment

Mar 23, 2026

Automatic Metrics Pairwise Preference Accuracy Not reported
RARE: Decoupling Representation Steering from Expert Routing in Mixture-of-Experts Language Models

Aug 21, 2026

Automatic Metrics Not reported Accuracy, Success rate Not reported
Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization

Aug 20, 2026

Automatic Metrics Not reported Accuracy Not reported
Decoupled Contrastive Decoding via Expert-Aligned Drafting

Aug 13, 2026

Not reported Not reported Latency Not reported
Researcher Workflow (Detailed)

Checklist

  • Gap: Human feedback

    Human feedback is present in 12 of 60 papers.

  • Gap: Quality controls

    Quality controls is present in 3 of 60 papers.

  • Strong: Benchmarks

    Benchmarks is present in 60 of 60 papers.

  • Strong: Metrics

    Metrics is present in 37 of 60 papers.

  • Gap: Known rater population

    Known rater population is present in 8 of 60 papers.

  • Gap: Known annotation unit

    Known annotation unit is present in 8 of 60 papers.

Strengths

  • Benchmarks is present in 60 of 60 papers.
  • Metrics is present in 37 of 60 papers.
  • Agentic evaluation is present in 9 of 60 papers.

Known Gaps

  • Human feedback is present in 12 of 60 papers.
  • Quality controls is present in 3 of 60 papers.
  • Known rater population is present in 8 of 60 papers.

Suggested Next Analyses

  • Review the most recent AIME or AlpacaEval or MMLU papers first, then compare reported metrics and quality-control context before treating results as comparable.

Recommended Queries

Known Limitations
  • This synthetic persisted page is generated from extraction data because the cached benchmark payload was missing for either-aime-or-alpacaeval-or-mmlu.
Research Utility Snapshot (Detailed)

Evaluation Modes

  • Automatic Metrics (33)
  • Llm As Judge (3)
  • Simulation Env (3)

Human Feedback Mix

  • None (48)
  • Pairwise Preference (8)
  • Demonstrations (2)
  • Expert Verification (1)

Top Benchmarks

  • MMLU (39)
  • AIME (17)
  • GSM8K (12)
  • MATH-500 (6)

Top Metrics

  • Accuracy (20)
  • Cost (8)
  • Latency (4)
  • Perplexity (3)

Top Papers On This Benchmark

Related Benchmark Hubs