Your ViT is Secretly an Image Segmentation Model
Fuente:
arXiv
Saved in:
| Main Authors: | Kerssies, Tommie, Cavagnero, Niccolò, Hermans, Alexander, Norouzi, Narges, Averta, Giuseppe, Leibe, Bastian, Dubbelman, Gijs, de Geus, Daan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VidEoMT: Your ViT is Secretly Also a Video Segmentation Model
by: Norouzi, Narges, et al.
Published: (2026)
by: Norouzi, Narges, et al.
Published: (2026)
PMT: Plain Mask Transformer for Image and Video Segmentation with Frozen Vision Encoders
by: Cavagnero, Niccolò, et al.
Published: (2026)
by: Cavagnero, Niccolò, et al.
Published: (2026)
How to Benchmark Vision Foundation Models for Semantic Segmentation?
by: Kerssies, Tommie, et al.
Published: (2024)
by: Kerssies, Tommie, et al.
Published: (2024)
First Place Solution to the ECCV 2024 BRAVO Challenge: Evaluating Robustness of Vision Foundation Models for Semantic Segmentation
by: Kerssies, Tommie, et al.
Published: (2024)
by: Kerssies, Tommie, et al.
Published: (2024)
ALGM: Adaptive Local-then-Global Token Merging for Efficient Semantic Segmentation with Plain Vision Transformers
by: Norouzi, Narges, et al.
Published: (2024)
by: Norouzi, Narges, et al.
Published: (2024)
Exploring the Benefits of Vision Foundation Models for Unsupervised Domain Adaptation
by: Englert, Brunó B., et al.
Published: (2024)
by: Englert, Brunó B., et al.
Published: (2024)
Task-aligned Part-aware Panoptic Segmentation through Joint Object-Part Representations
by: de Geus, Daan, et al.
Published: (2024)
by: de Geus, Daan, et al.
Published: (2024)
Towards Data-Efficient Video Pre-training with Frozen Image Foundation Models
by: Orlova, Svetlana, et al.
Published: (2026)
by: Orlova, Svetlana, et al.
Published: (2026)
What is the Added Value of UDA in the VFM Era?
by: Englert, Brunó B., et al.
Published: (2025)
by: Englert, Brunó B., et al.
Published: (2025)
Orion-Lite: Distilling LLM Reasoning into Efficient Vision-Only Driving Models
by: Gu, Jing, et al.
Published: (2026)
by: Gu, Jing, et al.
Published: (2026)
Simplifying Traffic Anomaly Detection with Video Foundation Models
by: Orlova, Svetlana, et al.
Published: (2025)
by: Orlova, Svetlana, et al.
Published: (2025)
A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens
by: Kerssies, Tommie, et al.
Published: (2026)
by: Kerssies, Tommie, et al.
Published: (2026)
Sa2VA-i: Improving Sa2VA Results with Consistent Training and Inference
by: Nekrasov, Alexey, et al.
Published: (2025)
by: Nekrasov, Alexey, et al.
Published: (2025)
How Important are Videos for Training Video LLMs?
by: Lydakis, George, et al.
Published: (2025)
by: Lydakis, George, et al.
Published: (2025)
DONUT: A Decoder-Only Model for Trajectory Prediction
by: Knoche, Markus, et al.
Published: (2025)
by: Knoche, Markus, et al.
Published: (2025)
Fine-Tuning Image-Conditional Diffusion Models is Easier than You Think
by: Garcia, Gonzalo Martin, et al.
Published: (2024)
by: Garcia, Gonzalo Martin, et al.
Published: (2024)
DINO in the Room: Leveraging 2D Foundation Models for 3D Segmentation
by: Knaebel, Karim, et al.
Published: (2025)
by: Knaebel, Karim, et al.
Published: (2025)
The revenge of BiSeNet: Efficient Multi-Task Image Segmentation
by: Rosi, Gabriele, et al.
Published: (2024)
by: Rosi, Gabriele, et al.
Published: (2024)
PEM: Prototype-based Efficient MaskFormer for Image Segmentation
by: Cavagnero, Niccolò, et al.
Published: (2024)
by: Cavagnero, Niccolò, et al.
Published: (2024)
Volume Transformer: Revisiting Vanilla Transformers for 3D Scene Understanding
by: Yilmaz, Kadir, et al.
Published: (2026)
by: Yilmaz, Kadir, et al.
Published: (2026)
Transient Fault Tolerant Semantic Segmentation for Autonomous Driving
by: Iurada, Leonardo, et al.
Published: (2024)
by: Iurada, Leonardo, et al.
Published: (2024)
Point2Vec for Self-Supervised Representation Learning on Point Clouds
by: Knaebel, Karim, et al.
Published: (2023)
by: Knaebel, Karim, et al.
Published: (2023)
OoDIS: Anomaly Instance Segmentation and Detection Benchmark
by: Nekrasov, Alexey, et al.
Published: (2024)
by: Nekrasov, Alexey, et al.
Published: (2024)
OpenSplat3D: Open-Vocabulary 3D Instance Segmentation using Gaussian Splatting
by: Piekenbrinck, Jens, et al.
Published: (2025)
by: Piekenbrinck, Jens, et al.
Published: (2025)
The BRAVO Semantic Segmentation Challenge Results in UNCV2024
by: Vu, Tuan-Hung, et al.
Published: (2024)
by: Vu, Tuan-Hung, et al.
Published: (2024)
VFM-UDA++: Improving Network Architectures and Data Strategies for Unsupervised Domain Adaptive Semantic Segmentation
by: Englert, Brunó B., et al.
Published: (2025)
by: Englert, Brunó B., et al.
Published: (2025)
SurGe: Improved Surface Geometry in Point Maps
by: Knaebel, Karim, et al.
Published: (2026)
by: Knaebel, Karim, et al.
Published: (2026)
Vanilla ViT for Automotive Point Cloud Semantic Segmentation
by: Puy, Gilles, et al.
Published: (2026)
by: Puy, Gilles, et al.
Published: (2026)
Applying ViT in Generalized Few-shot Semantic Segmentation
by: Geng, Liyuan, et al.
Published: (2024)
by: Geng, Liyuan, et al.
Published: (2024)
Mask4Former: Mask Transformer for 4D Panoptic Segmentation
by: Yilmaz, Kadir, et al.
Published: (2023)
by: Yilmaz, Kadir, et al.
Published: (2023)
Point-VOS: Pointing Up Video Object Segmentation
by: Zulfikar, Idil Esen, et al.
Published: (2024)
by: Zulfikar, Idil Esen, et al.
Published: (2024)
DeNAS-ViT: Data Efficient NAS-Optimized Vision Transformer for Ultrasound Image Segmentation
by: Chen, Renqi, et al.
Published: (2024)
by: Chen, Renqi, et al.
Published: (2024)
Searching on a Budget: HW-NAS with 10 Latency Probes
by: Capuano, Francesco, et al.
Published: (2025)
by: Capuano, Francesco, et al.
Published: (2025)
Deeper Inside Deep ViT
by: Hong, Sungrae
Published: (2025)
by: Hong, Sungrae
Published: (2025)
An Ordinal Regression Framework for a Deep Learning Based Severity Assessment for Chest Radiographs
by: Wienholt, Patrick, et al.
Published: (2024)
by: Wienholt, Patrick, et al.
Published: (2024)
SPAR: Single-Pass Any-Resolution ViT for Open-vocabulary Segmentation
by: Kombol, Naomi, et al.
Published: (2026)
by: Kombol, Naomi, et al.
Published: (2026)
I&S-ViT: An Inclusive & Stable Method for Pushing the Limit of Post-Training ViTs Quantization
by: Zhong, Yunshan, et al.
Published: (2023)
by: Zhong, Yunshan, et al.
Published: (2023)
SANSA: Unleashing the Hidden Semantics in SAM2 for Few-Shot Segmentation
by: Cuttano, Claudia, et al.
Published: (2025)
by: Cuttano, Claudia, et al.
Published: (2025)
Look Gauss, No Pose: Novel View Synthesis using Gaussian Splatting without Accurate Pose Initialization
by: Schmidt, Christian, et al.
Published: (2024)
by: Schmidt, Christian, et al.
Published: (2024)
MaskTerial: A Foundation Model for Automated 2D Material Flake Detection
by: Uslu, Jan-Lucas, et al.
Published: (2024)
by: Uslu, Jan-Lucas, et al.
Published: (2024)
Similar Items
-
VidEoMT: Your ViT is Secretly Also a Video Segmentation Model
by: Norouzi, Narges, et al.
Published: (2026) -
PMT: Plain Mask Transformer for Image and Video Segmentation with Frozen Vision Encoders
by: Cavagnero, Niccolò, et al.
Published: (2026) -
How to Benchmark Vision Foundation Models for Semantic Segmentation?
by: Kerssies, Tommie, et al.
Published: (2024) -
First Place Solution to the ECCV 2024 BRAVO Challenge: Evaluating Robustness of Vision Foundation Models for Semantic Segmentation
by: Kerssies, Tommie, et al.
Published: (2024) -
ALGM: Adaptive Local-then-Global Token Merging for Efficient Semantic Segmentation with Plain Vision Transformers
by: Norouzi, Narges, et al.
Published: (2024)