OmAgent: A Multi-modal Agent Framework for Complex Video Understanding with Task Divide-and-Conquer
Results and benchmarks
OmAgent: A Multi-modal Agent Framework for Complex Video Understanding with Task Divide-and-Conquer focuses on agentic tool use.
Benchmark evidence is limited
Evidence graph: 3 refs, 3 links.
Utility signals: depth 60/100, grounding 75/100, status medium.
Implementation
Best maintained implementation now
[EMNLP-2024] Build multimodal language agents for fast prototype and production
2,665 stars · 292 forks · Last push Mar 19, 2025 · Apache-2.0 license
- License
- CI
- Dependencies
- Docker
Official implementation from Papers with Code · Repository link is mentioned in the paper metadata · Matched via arXiv identifier search
om-ai-lab/OmAgent is the strongest maintained implementation based on ranking signals. CI workflows are present. License is declared (Apache-2.0).
Open om-ai-lab/OmAgent- Dependency manifest is missing
- Selected om-ai-lab/OmAgent as the strongest maintained implementation for new work.
- Includes CI workflow signals.
- Repository activity is within the last 24 months.
Compare implementation paths
Compare maintenance quality, reproducibility coverage, and evidence confidence before choosing a reproduction baseline.
- Maintenance
- Stale
- Confidence
- High
- Reproducibility
- Moderate
- Stars
- 2,665
- Last push
- Mar 19, 2025 (524d)
Official implementation from Papers with Code · Repository link is mentioned in the paper metadata
- No push in 12+ months
- No Docker setup
- Dependency manifest missing
- Maintenance
- Stale risk
- Confidence
- Low
- Reproducibility
- Limited
- Stars
- 91
- Last push
- Nov 20, 2025 (279d)
Community adoption signal (91 stars)
- No CI pipeline detected
- No tagged releases
- No Docker setup
- Maintenance
- Stale
- Confidence
- Low
- Reproducibility
- Strong
- Stars
- 0
- Last push
- Dec 24, 2024 (610d)
Matched via arXiv identifier search
- No push in 12+ months
- No tagged releases
- No Docker setup
Reproduction readiness
Major work
No dependency manifest, manual reconstruction required
- om-ai-lab/OmAgent has no requirements.txt, environment.yml, pyproject.toml, or Dockerfile.
- You will need to reverse-engineer dependencies from import statements in the source code.
- Last push was 524 days ago.
Hardware requirements
- Expect multi-day setup/compute for meaningful reproduction based on current guidance.
Validation caveat
Repositories and ecosystem
No additional verified repositories beyond the primary recommendation.
These repositories had low-confidence matching signals and are hidden by default.
- om-ai-lab/ZoomEye
Confidence: Low · 91 stars
- Silentharry94/justdoit
Confidence: Low · 0 stars
- AnotiusRebel/OmAgent
Confidence: Low · 0 stars
Hugging Face artifacts
No trustworthy direct or curated related Hugging Face artifacts were found yet. Use targeted searches to quickly locate candidate models, datasets, and demos.
Tip: start with models, then check datasets and spaces if you need evaluation data or demos.
Research context
Tasks
Agentic tool use, Video understanding / reasoning
Methods
Agentic systems
Domains
Computer vision, AI Agents
Open this paper in HFEPX to review benchmark signals, evaluation modes, and human-feedback protocol context.
Open in HFEPXJump to Paper2Code search queries derived from this paper's research context.
Data includes links from Papers with Code ( CC-BY-SA-4.0 ).