Vision Transformers: From Semantic Segmentation to Dense Prediction
Results and benchmarks
Vision Transformers: From Semantic Segmentation to Dense Prediction presents a transformer approach for image classification.
| Task | Dataset | Metric | Value | Source |
|---|---|---|---|---|
| Image classification | FCN (160k, SS) OpenMMLab 2020 | mIoU. | 39.9 | paper-derived |
| Image classification | FCN (40k, SS) OpenMMLab 2020 | mIoU. | 73.9 | paper-derived |
Audit each benchmark finding before selecting an implementation path. Evidence refs map to the disclosure below.
Evidence graph: 4 refs, 4 links.
Utility signals: depth 100/100, grounding 95/100, status high.
Implementation
No direct paper-linked artifacts found; showing strongest related artifacts
Maintained implementation evidence is not confirmed for this paper yet.
Use the implementation status and reproduction sections for the current action plan.
No verified maintained repo yet
There is no verified maintained implementation yet. Use this baseline plan to decide whether to prototype now or defer.
- No maintained paper-verified implementation was found; start with the closest related repositories below.
- Compare repo methods against the paper equations/algorithm before trusting metrics.
- Create a minimal baseline implementation from the paper and use adjacent repos as references.
Time to first repro: a few days · Best available artifact: facebook/mask2former-swin-large-ade-semantic
WangLibo1995/GeoSeg is the closest maintained adjacent implementation (Matches contextual method/domain keyword: transformer). It is not paper-verified; validate algorithm and evaluation setup against the paper before trusting reported metrics. Community adoption signal: 1096 GitHub stars.
- Adjacent implementations are not paper-verified
- Recommended repository is adjacent and not paper-verified.
- Adjacent implementation match confidence is low.
Reproduction readiness
No repo
No verified implementation available
- No maintained repository has been identified for this paper. Check adjacent implementations or HF artifacts below.
Hardware requirements
- Expect multi-day setup/compute for meaningful reproduction based on current guidance.
Framework baselines
- Hugging Face Transformers training guide
Modern transformer training baseline.
- PyTorch nn.Transformer docs
Reference transformer building block implementation.
Repositories and ecosystem
Closest related implementations
These are not paper-verified. Use them as reference points when no direct implementation is available.
- WangLibo1995/GeoSeg Adjacent · Confidence: Low · 1,096 stars
Matches contextual method/domain keyword: transformer
- zhangyp15/OccFormer Adjacent · Confidence: Low · 411 stars
Matches contextual method/domain keyword: transformer
- zbwxp/SegVit Adjacent · Confidence: Low · 263 stars
Matches contextual method/domain keyword: transformer
- facebookresearch/HRViT Adjacent · Confidence: Low · 199 stars
Matches contextual method/domain keyword: transformer
No additional verified repositories beyond the primary recommendation.
Hugging Face artifacts
No direct paper-linked artifacts were found. Showing strongest curated related artifacts for faster exploration.
Models
- facebook/mask2former-swin-large-ade-semantic
387,743 downloads · 24 likes
- facebook/mask2former-swin-large-cityscapes-semantic
173,052 downloads · 38 likes
- facebook/mask2former-swin-large-mapillary-vistas-semantic
66,413 downloads · 9 likes
Broaden model search
Datasets
No trustworthy datasets matches right now.
Search datasets on Hugging FaceSpaces
No trustworthy spaces matches right now.
Search spaces on Hugging FaceResearch context
Tasks
Image classification
Methods
Transformer
Domains
Computer vision
Open this paper in HFEPX to review benchmark signals, evaluation modes, and human-feedback protocol context.
Open in HFEPXJump to Paper2Code search queries derived from this paper's research context.