The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants
Results and benchmarks
The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants is the primary contribution described in this paper.
Benchmark evidence is limited
Evidence graph: 3 refs, 3 links.
Utility signals: depth 65/100, grounding 75/100, status medium.
Implementation
Best maintained implementation now
Repo for the Belebele dataset, a massively multilingual reading comprehension dataset.
339 stars · 20 forks · Last push Dec 18, 2024 · NOASSERTION license
- License
- CI
- Dependencies
- Docker
Official implementation from Papers with Code · Strong overlap with paper title keywords · Community adoption signal (339 stars)
facebookresearch/belebele is the strongest maintained implementation based on ranking signals. License is declared (NOASSERTION).
Open facebookresearch/belebele- No CI workflows detected
- Dependency manifest is missing
- Selected facebookresearch/belebele as the strongest maintained implementation for new work.
- Repository activity is within the last 24 months.
Compare implementation paths
Compare maintenance quality, reproducibility coverage, and evidence confidence before choosing a reproduction baseline.
- Maintenance
- Stale
- Confidence
- High
- Reproducibility
- Limited
- Stars
- 339
- Last push
- Dec 18, 2024 (616d)
Official implementation from Papers with Code · Strong overlap with paper title keywords
- No push in 12+ months
- No CI pipeline detected
- No tagged releases
- Maintenance
- Active
- Confidence
- Low
- Reproducibility
- Strong
- Stars
- 16
- Last push
- Aug 24, 2026 (2d)
Matched via arXiv identifier search
- No Docker setup
- Low confidence match
- Maintenance
- Stale risk
- Confidence
- Low
- Reproducibility
- Limited
- Stars
- 23
- Last push
- Jan 17, 2026 (221d)
Matched via arXiv identifier search
- No CI pipeline detected
- No tagged releases
- No Docker setup
Reproduction readiness
Major work
No dependency manifest, manual reconstruction required
- facebookresearch/belebele has no requirements.txt, environment.yml, pyproject.toml, or Dockerfile.
- You will need to reverse-engineer dependencies from import statements in the source code.
- Last push was 616 days ago.
Hardware requirements
- Expect multi-day setup/compute for meaningful reproduction based on current guidance.
Validation caveat
Repositories and ecosystem
No additional verified repositories beyond the primary recommendation.
These repositories had low-confidence matching signals and are hidden by default.
- ModelCloud/Evalution
Confidence: Low · 16 stars
- jphall663/gai_risk_management
Confidence: Low · 23 stars
- isShayulajiao/CCL25-Eval-ZhengMing
Confidence: Low · 10 stars
- Pseudo-Lab/Korean_LLM_Benchmark_Test
Confidence: Low · 5 stars
Hugging Face artifacts
No trustworthy direct or curated related Hugging Face artifacts were found yet. Use targeted searches to quickly locate candidate models, datasets, and demos.
Tip: start with models, then check datasets and spaces if you need evaluation data or demos.
Research context
Open this paper in HFEPX to review benchmark signals, evaluation modes, and human-feedback protocol context.
Open in HFEPXData includes links from Papers with Code ( CC-BY-SA-4.0 ).