Region-Aware Pretraining for Open-Vocabulary Object Detection with Vision Transformers
Results and benchmarks
Region-Aware Pretraining for Open-Vocabulary Object Detection with Vision Transformers presents a transformer approach for image classification.
| Task | Dataset | Metric | Value | Source |
|---|---|---|---|---|
| Detection | COCO | AP | 50 | paper-derived |
| Image classification | Objects365 | AP | 50 | paper-derived |
| Language modeling | DetPro-Cascade du2022learning | AP. | 27.0 | paper-derived |
| Language modeling | ViLD gu2022openvocabulary | AP. | 51.3 | paper-derived |
Audit each benchmark finding before selecting an implementation path. Evidence refs map to the disclosure below.
Evidence graph: 4 refs, 4 links.
Utility signals: depth 100/100, grounding 95/100, status high.
Implementation
Best maintained implementation now
Google Research
38,611 stars · 8,464 forks · Last push Aug 21, 2026 · Apache-2.0 license
- License
- CI
- Dependencies
- Docker
Official implementation from Papers with Code · Community adoption signal (38611 stars)
google-research/google-research is the strongest maintained implementation based on ranking signals. License is declared (Apache-2.0).
Open google-research/google-research- No CI workflows detected
- Dependency manifest is missing
- Selected google-research/google-research as the strongest maintained implementation for new work.
- Repository activity is within the last 24 months.
- Official repository is preserved separately as historical context.
Compare implementation paths
Compare maintenance quality, reproducibility coverage, and evidence confidence before choosing a reproduction baseline.
- Maintenance
- Active
- Confidence
- High
- Reproducibility
- Limited
- Stars
- 38,611
- Last push
- Aug 21, 2026 (4d)
Official implementation from Papers with Code · Community adoption signal (38611 stars)
- No CI pipeline detected
- No tagged releases
- No Docker setup
- Maintenance
- Stale
- Confidence
- High
- Reproducibility
- Limited
- Stars
- 17
- Last push
- Aug 24, 2023 (1097d)
Official implementation from Papers with Code · Repository link is mentioned in the paper metadata
- No push in 12+ months
- No CI pipeline detected
- No tagged releases
Reproduction readiness
Major work
No dependency manifest, manual reconstruction required
- google-research/google-research has no requirements.txt, environment.yml, pyproject.toml, or Dockerfile.
- You will need to reverse-engineer dependencies from import statements in the source code.
Hardware requirements
- Expect multi-day setup/compute for meaningful reproduction based on current guidance.
Hugging Face artifacts
No direct paper-linked artifacts were found. Showing strongest curated related artifacts for faster exploration.
Models
- ShuaiYang03/instructvla_pretraining_v2_libero_goal_wrist-image_aug
34 downloads · 1 likes
- JustinAngel/workshop-v1-pretraining
885 downloads · 5 likes
- FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
594 downloads · 2 likes
Broaden model search
Datasets
- applied-ai-018/pretraining_v1-omega_books
663,150 downloads · 22 likes · Updated Aug 5, 2024
- nvidia/Nemotron-Pretraining-Code-v2
3,791 downloads · 130 likes · Updated Dec 22, 2025
Spaces
Research context
Tasks
Image classification
Methods
Transformer
Domains
Computer vision
Open this paper in HFEPX to review benchmark signals, evaluation modes, and human-feedback protocol context.
Open in HFEPXJump to Paper2Code search queries derived from this paper's research context.
Data includes links from Papers with Code ( CC-BY-SA-4.0 ).