HI-MoE: Hierarchical Instance-Conditioned Mixture-of-Experts for Object Detection

Q: How reproducible is "HI-MoE: Hierarchical Instance-Conditioned Mixture-of-Experts for Object Detection"?

Estimated time to first reproduction: a few days. Risk flags: No repository-level reproducibility signals are currently available, Estimate is based on paper-only reproduction flow. No direct maintained implementation was found. Use the paper PDF and citation graph to design a baseline reproduction.

Q: What framework is used to implement "HI-MoE: Hierarchical Instance-Conditioned Mixture-of-Experts for Object Detection"?

The primary implementation uses TorchVision object detection finetuning tutorial.

Vadim Vashkelis, Natalia Trukhina

Published: Apr 6, 2026

No direct implementation yet

Evidence: Inferred

Domain fit: AI-adjacent

Verified repos: 0

Paper appears method- or tooling-adjacent to AI workflows with partial ecosystem coverage.

Framework: TorchVision object detection finetuning tutorial

Time to first repro: a few days

2 risk flags

arXiv PDF

Mixture-of-Experts (MoE) architectures enable conditional computation by activating only a subset of model parameters for each input. Although sparse routing has been highly effective in language models and has also shown promise in vision, most vision MoE methods operate at the image or patch level. This granularity is poorly aligned with object detection, where the fundamental unit of reasoning is an object query c ...

Read full abstract

orresponding to a candidate instance. We propose Hierarchical Instance-Conditioned Mixture-of-Experts (HI-MoE), a DETR-style detection architecture that performs routing in two stages: a lightweight scene router first selects a scene-consistent expert subset, and an instance router then assigns each object query to a small number of experts within that subset. This design aims to preserve sparse computation while better matching the heterogeneous, instance-centric structure of detection. In the current draft, experiments are concentrated on COCO with preliminary specialization analysis on LVIS. Under these settings, HI-MoE improves over a dense DINO baseline and over simpler token-level or instance-only routing variants, with especially strong gains on small objects. We also provide an initial visualization of expert specialization patterns. We present the method, ablations, and current limitations in a form intended to support further experimental validation.

Technical details

Canonical key: arxiv-2604.04908

Cache status: Stale (SWR served)

Generated at: Apr 26, 2026, 9:52 PM

Artifact coverage: sparse

HF provider: ok (token)

PWC source used: No

LLM status: not_generated

LLM model: n/a

LLM generated: Unknown

LLM content type: n/a

HF policy: hf-relevance-v27

context only

Benchmarks: thin evidence

Time to repro: a few days

2 risk flags

TorchVision object detection finetuning tutorial

Results & Benchmarks

Freshness tier: hot

Direct + Inferred Evidence

Computer vision

COCO

280

Source: paper fulltext

Computer vision

DETR

42.0

Source: paper fulltext

Computer vision

Deformable DETR

48.7

Source: paper fulltext

Computer vision

DINO

51.3

Source: paper fulltext

Benchmark evidence drill-down

4 findings

Audit each benchmark finding before selecting an implementation path. Evidence refs map to the disclosure section below.

Task	Dataset	Metric	Value	Source	Evidence refs
Computer vision	COCO	AP	280	paper-derived	No explicit refs
Computer vision	DETR	AP	42.0	paper-derived	No explicit refs
Computer vision	Deformable DETR	AP	48.7	paper-derived	No explicit refs
Computer vision	DINO	AP	51.3	paper-derived	No explicit refs

Mixture-of-Experts (MoE) architectures enable conditional computation by activating only a subset of model parameters for each input.

Implementation Evidence Summary

Confidence: low

Recommendation evidence is currently too limited for a maintained-repo choice. Use Implementation Status and Reproduction Path for a practical baseline plan.

Reproduction Risks

Estimate is based on paper-only reproduction flow

Hardware Notes

Expect multi-day setup/compute for meaningful reproduction based on current guidance.

Evidence disclosure

Evidence graph: 2 refs, 1 links.

Utility signals: depth 95/100, grounding 68/100, status medium.

Implementation Status

No verified maintained repo

There is no verified maintained implementation yet. Use this baseline plan to decide whether to prototype now or defer.

No direct maintained implementation was found. Use the paper PDF and citation graph to design a baseline reproduction.
Track assumptions and missing details in an experiment log before coding.

Time to first repro: a few days

Reproduction readiness

No Repo

Time to first repro: days

Last checked: Apr 26, 2026

Hardware requirements

Expect multi-day setup/compute for meaningful reproduction based on current guidance.

No verified implementation available

· No maintained repository has been identified for this paper. Check adjacent implementations or HF artifacts below.

Framework baselines

TorchVision object detection finetuning tutorial
Baseline setup for object detection workflows.

Hugging Face artifacts

No trustworthy direct or curated related Hugging Face artifacts were found yet.

Continue with targeted Hugging Face searches derived from the paper title and method context:

Models

arxiv:2604.04908 HI-MoE Instance-Conditioned

Datasets

arxiv:2604.04908 HI-MoE dataset

Spaces

arxiv:2604.04908 HI-MoE demo

Tip: start with models, then check datasets/spaces if you need evaluation data or demos.

Direct artifact matches are currently sparse. Use targeted Hugging Face searches to quickly locate candidate models, datasets, and demos.

Search models Search datasets Search spaces

Research context

Tasks

Computer vision

Methods

Transformer

Domains

Computer vision, Natural Language Processing

Evaluation & Human Feedback Data

Open this paper in HFEPX to review benchmark signals, evaluation modes, and human-feedback protocol context.

Open in HFEPX

Explore Similar Papers

Jump to Paper2Code search queries derived from this paper's research context.

Computer vision Transformer Natural Language Processing

Need human evaluators for your AI research? Scale annotation with expert AI Trainers.

Post a Job Get a Quote