MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
Results and benchmarks
MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI focuses on reasoning / puzzle solving.
Benchmark evidence is limited
Evidence graph: 4 refs, 4 links.
Utility signals: depth 65/100, grounding 85/100, status medium.
Implementation
Best maintained implementation now
This repo contains evaluation code for the paper "MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI"
592 stars · 56 forks · Last push Jul 28, 2026 · Apache-2.0 license
- License
- CI
- Dependencies
- Docker
Official implementation from Papers with Code · Strong overlap with paper title keywords · Community adoption signal (592 stars)
MMMU-Benchmark/MMMU is the strongest maintained implementation based on ranking signals. License is declared (Apache-2.0).
Open MMMU-Benchmark/MMMU- No CI workflows detected
- Dependency manifest is missing
- Selected MMMU-Benchmark/MMMU as the strongest maintained implementation for new work.
- Repository activity is within the last 24 months.
Compare implementation paths
Compare maintenance quality, reproducibility coverage, and evidence confidence before choosing a reproduction baseline.
- Maintenance
- Active
- Confidence
- High
- Reproducibility
- Limited
- Stars
- 592
- Last push
- Jul 28, 2026 (30d)
Official implementation from Papers with Code · Strong overlap with paper title keywords
- No CI pipeline detected
- No tagged releases
- No Docker setup
- Maintenance
- Stale
- Confidence
- Low
- Reproducibility
- Strong
- Stars
- 7,836
- Last push
- Nov 27, 2024 (637d)
Community adoption signal (7836 stars)
- No push in 12+ months
- No tagged releases
- Low confidence match
- Maintenance
- Recently updated
- Confidence
- Low
- Reproducibility
- Limited
- Stars
- 25
- Last push
- May 12, 2026 (106d)
Community adoption signal (25 stars)
- No CI pipeline detected
- No tagged releases
- No Docker setup
Reproduction readiness
Major work
No dependency manifest, manual reconstruction required
- MMMU-Benchmark/MMMU has no requirements.txt, environment.yml, pyproject.toml, or Dockerfile.
- You will need to reverse-engineer dependencies from import statements in the source code.
Hardware requirements
- Expect multi-day setup/compute for meaningful reproduction based on current guidance.
Validation caveat
Repositories and ecosystem
No additional verified repositories beyond the primary recommendation.
These repositories had low-confidence matching signals and are hidden by default.
- 01-ai/yi
Confidence: Low · 7,836 stars
- eric-ai-lab/probmed
Confidence: Low · 25 stars
- wjlgatech/FM-os
Confidence: Low · 3 stars
- VyetGokyra/awaresome_LLM_eval_benchmark
Confidence: Low · 8 stars
Hugging Face artifacts
No direct paper-linked artifacts were found. Showing strongest curated related artifacts for faster exploration.
Models
- AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS
32,784 downloads · 59 likes
- AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-NVFP4-MTP
10,030 downloads · 37 likes
Broaden model search
Datasets
- Perle-ai/multimodal-ct-radiology-reports
423 downloads · 1 likes · Updated May 6, 2026
- Voxel51/fomo-multimodal-sample
392 downloads · 1 likes · Updated Aug 7, 2026
Broaden dataset search
Spaces
No trustworthy spaces matches right now.
Search spaces on Hugging FaceResearch context
Tasks
Reasoning / puzzle solving
Methods
None detected
Domains
None detected
Open this paper in HFEPX to review benchmark signals, evaluation modes, and human-feedback protocol context.
Open in HFEPXJump to Paper2Code search queries derived from this paper's research context.
Data includes links from Papers with Code ( CC-BY-SA-4.0 ).