Results and benchmarks
Measuring Coding Challenge Competence With APPS is the primary contribution described in this paper.
Benchmark evidence is limited
Evidence graph: 3 refs, 3 links.
Utility signals: depth 100/100, grounding 85/100, status high.
Implementation
Historical official implementation (not recommended for new builds)
Only a historical official implementation is available
Use with caution for new projects; verify against current tooling and maintained community alternatives.
hendrycks/apps · 538 stars · Last push Jun 19, 2024
codedotal/gpt-code-clippy is the closest maintained adjacent implementation (Community adoption signal (3264 stars)). It is not paper-verified; validate algorithm and evaluation setup against the paper before trusting reported metrics. Community adoption signal: 3264 GitHub stars.
Open hendrycks/apps- Adjacent implementations are not paper-verified
- Recommended repository is adjacent and not paper-verified.
- Adjacent implementation match confidence is low.
- No direct maintained implementation is currently verified.
- Only historical official repository was found: hendrycks/apps.
- No maintained paper-verified implementation met reliability thresholds.
Compare implementation paths
Compare maintenance quality, reproducibility coverage, and evidence confidence before choosing a reproduction baseline.
- Maintenance
- Stale
- Confidence
- High
- Reproducibility
- Moderate
- Stars
- 538
- Last push
- Jun 19, 2024 (798d)
Official implementation from Papers with Code · Repository link is mentioned in the paper metadata
- No push in 12+ months
- No CI pipeline detected
- No tagged releases
- Maintenance
- Stale
- Confidence
- Low
- Reproducibility
- Limited
- Stars
- 8
- Last push
- Aug 12, 2025 (380d)
Matched via arXiv identifier search
- No push in 12+ months
- No CI pipeline detected
- No tagged releases
- Maintenance
- Stale
- Confidence
- Low
- Reproducibility
- Limited
- Stars
- 1
- Last push
- Jul 21, 2025 (401d)
Matched via arXiv identifier search
- No push in 12+ months
- No CI pipeline detected
- No tagged releases
Reproduction readiness
Setup required
Dependencies pinned, manual setup needed
- hendrycks/apps has requirements.txt but requires manual environment setup.
- Last push was 798 days ago, so expect possible dependency version conflicts.
- No Dockerfile, so you will set up the environment manually.
- No CI pipeline, so test coverage is unknown.
Hardware requirements
- Expect multi-day setup/compute for meaningful reproduction based on current guidance.
Quick start
git clone https://github.com/hendrycks/apps.git
pip install -r requirements.txt Repositories and ecosystem
Closest related implementations
These are not paper-verified. Use them as reference points when no direct implementation is available.
- codedotal/gpt-code-clippy Adjacent · Confidence: Low · 3,264 stars
Community adoption signal (3264 stars)
- ncoop57/gpt-code-clippy Adjacent · Confidence: Low · 3,264 stars
Community adoption signal (3264 stars)
No additional verified repositories beyond the primary recommendation.
These repositories had low-confidence matching signals and are hidden by default.
- VyetGokyra/awaresome_LLM_eval_benchmark
Confidence: Low · 8 stars
- juzizi44/GRPO4CodeGen
Confidence: Low · 1 stars
- HyunjoonCho/mvce
Confidence: Low · 0 stars
Hugging Face artifacts
No trustworthy direct or curated related Hugging Face artifacts were found yet. Use targeted searches to quickly locate candidate models, datasets, and demos.
Tip: start with models, then check datasets and spaces if you need evaluation data or demos.
Research context
Open this paper in HFEPX to review benchmark signals, evaluation modes, and human-feedback protocol context.
Open in HFEPXData includes links from Papers with Code ( CC-BY-SA-4.0 ).