Skip to content
OpenTrain AIFor AI Companies
← Back to explorer

HFEPX Benchmark Hub

AIME or HumanEval+ or MMLU Benchmark Papers

Snapshot fallback from 2026-08-31. This benchmark page remains available while the live HFEPX payload refreshes; use the full paper list after the API recovers.

Papers: 57 Last published: Mar 28, 2026 Global RSS

Researcher Quick Triage

Use this page for benchmark-matched method comparisons and eval protocol selection. Quality band: Developing .

Analysis blocks are computed from the loaded sample (0 of 57 papers).

High-Signal Coverage

0%

0 / 0 sampled papers are not low-signal flagged.

Replication-Ready Set

0

Papers with explicit benchmark + metric + eval mode fields.

Quality Controls

0%

0 papers report calibration/adjudication/IAA controls.

  • 0 papers explicitly name benchmark datasets in the sampled set.
  • 0 papers report at least one metric term in metadata extraction.
  • Start with the ranked shortlist below before reading all papers.

Primary action: Use this page to map benchmark mentions first; wait for stronger metric/QC coverage before strict comparisons.

Why This Matters (Expanded)

Why This Matters For Eval Research

  • The route remains reachable for users and monitors instead of returning an upstream 503.
Protocol Notes (Expanded)

Protocol Takeaways

  • Paper-level protocol details are temporarily unavailable in the fallback response.

Benchmark Interpretation

  • AIME benchmark signal
  • HumanEval+ benchmark signal
  • MMLU benchmark signal
Researcher Workflow (Detailed)

Checklist

  • Gap: Human feedback

    Human feedback coverage requires the live HFEPX payload.

  • Gap: Quality controls

    Quality controls coverage requires the live HFEPX payload.

  • Gap: Benchmarks

    Benchmarks coverage requires the live HFEPX payload.

  • Gap: Metrics

    Metrics coverage requires the live HFEPX payload.

  • Gap: Known rater population

    Known rater population coverage requires the live HFEPX payload.

  • Gap: Known annotation unit

    Known annotation unit coverage requires the live HFEPX payload.

Known Gaps

  • Live paper-level details were unavailable during this request.

Suggested Next Analyses

  • Refresh this page after the HFEPX API recovers to inspect paper-level protocol details.

Recommended Queries

Known Limitations
  • This page is a snapshot fallback generated because the live benchmark hub API failed for this request.
Research Utility Snapshot (Detailed)

Evaluation Modes

Human Feedback Mix

Top Benchmarks

  • AIME (57)
  • HumanEval+ (57)
  • MMLU (57)

Top Metrics

Top Papers On This Benchmark

No papers available for this benchmark yet.

Related Benchmark Hubs