PMT: Plain Mask Transformer for Image and Video Segmentation with Frozen Vision Encoders
Fuente:
arXiv
Saved in:
| Main Authors: | Cavagnero, Niccolò, Norouzi, Narges, Dubbelman, Gijs, de Geus, Daan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ALGM: Adaptive Local-then-Global Token Merging for Efficient Semantic Segmentation with Plain Vision Transformers
by: Norouzi, Narges, et al.
Published: (2024)
by: Norouzi, Narges, et al.
Published: (2024)
VidEoMT: Your ViT is Secretly Also a Video Segmentation Model
by: Norouzi, Narges, et al.
Published: (2026)
by: Norouzi, Narges, et al.
Published: (2026)
Your ViT is Secretly an Image Segmentation Model
by: Kerssies, Tommie, et al.
Published: (2025)
by: Kerssies, Tommie, et al.
Published: (2025)
Towards Data-Efficient Video Pre-training with Frozen Image Foundation Models
by: Orlova, Svetlana, et al.
Published: (2026)
by: Orlova, Svetlana, et al.
Published: (2026)
Task-aligned Part-aware Panoptic Segmentation through Joint Object-Part Representations
by: de Geus, Daan, et al.
Published: (2024)
by: de Geus, Daan, et al.
Published: (2024)
Orion-Lite: Distilling LLM Reasoning into Efficient Vision-Only Driving Models
by: Gu, Jing, et al.
Published: (2026)
by: Gu, Jing, et al.
Published: (2026)
How to Benchmark Vision Foundation Models for Semantic Segmentation?
by: Kerssies, Tommie, et al.
Published: (2024)
by: Kerssies, Tommie, et al.
Published: (2024)
First Place Solution to the ECCV 2024 BRAVO Challenge: Evaluating Robustness of Vision Foundation Models for Semantic Segmentation
by: Kerssies, Tommie, et al.
Published: (2024)
by: Kerssies, Tommie, et al.
Published: (2024)
Exploring the Benefits of Vision Foundation Models for Unsupervised Domain Adaptation
by: Englert, Brunó B., et al.
Published: (2024)
by: Englert, Brunó B., et al.
Published: (2024)
A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens
by: Kerssies, Tommie, et al.
Published: (2026)
by: Kerssies, Tommie, et al.
Published: (2026)
PEM: Prototype-based Efficient MaskFormer for Image Segmentation
by: Cavagnero, Niccolò, et al.
Published: (2024)
by: Cavagnero, Niccolò, et al.
Published: (2024)
The revenge of BiSeNet: Efficient Multi-Task Image Segmentation
by: Rosi, Gabriele, et al.
Published: (2024)
by: Rosi, Gabriele, et al.
Published: (2024)
VFM-UDA++: Improving Network Architectures and Data Strategies for Unsupervised Domain Adaptive Semantic Segmentation
by: Englert, Brunó B., et al.
Published: (2025)
by: Englert, Brunó B., et al.
Published: (2025)
Simplifying Traffic Anomaly Detection with Video Foundation Models
by: Orlova, Svetlana, et al.
Published: (2025)
by: Orlova, Svetlana, et al.
Published: (2025)
How Important are Videos for Training Video LLMs?
by: Lydakis, George, et al.
Published: (2025)
by: Lydakis, George, et al.
Published: (2025)
Volume Transformer: Revisiting Vanilla Transformers for 3D Scene Understanding
by: Yilmaz, Kadir, et al.
Published: (2026)
by: Yilmaz, Kadir, et al.
Published: (2026)
DONUT: A Decoder-Only Model for Trajectory Prediction
by: Knoche, Markus, et al.
Published: (2025)
by: Knoche, Markus, et al.
Published: (2025)
What is the Added Value of UDA in the VFM Era?
by: Englert, Brunó B., et al.
Published: (2025)
by: Englert, Brunó B., et al.
Published: (2025)
DINO in the Room: Leveraging 2D Foundation Models for 3D Segmentation
by: Knaebel, Karim, et al.
Published: (2025)
by: Knaebel, Karim, et al.
Published: (2025)
Revisiting Radar Perception With Spectral Point Clouds
by: Alsharif, Hamza, et al.
Published: (2026)
by: Alsharif, Hamza, et al.
Published: (2026)
REFNet++: Multi-Task Efficient Fusion of Camera and Radar Sensor Data in Bird's-Eye Polar View
by: Chandrasekaran, Kavin, et al.
Published: (2026)
by: Chandrasekaran, Kavin, et al.
Published: (2026)
A Resource Efficient Fusion Network for Object Detection in Bird's-Eye View using Camera and Raw Radar Data
by: Chandrasekaran, Kavin, et al.
Published: (2024)
by: Chandrasekaran, Kavin, et al.
Published: (2024)
Fine-Tuning Image-Conditional Diffusion Models is Easier than You Think
by: Garcia, Gonzalo Martin, et al.
Published: (2024)
by: Garcia, Gonzalo Martin, et al.
Published: (2024)
Transient Fault Tolerant Semantic Segmentation for Autonomous Driving
by: Iurada, Leonardo, et al.
Published: (2024)
by: Iurada, Leonardo, et al.
Published: (2024)
HyperTopo-Adapters: Geometry- and Topology-Aware Segmentation of Leaf Lesions on Frozen Encoders
by: Ndubuisi, Chimdi Walter, et al.
Published: (2025)
by: Ndubuisi, Chimdi Walter, et al.
Published: (2025)
Sa2VA-i: Improving Sa2VA Results with Consistent Training and Inference
by: Nekrasov, Alexey, et al.
Published: (2025)
by: Nekrasov, Alexey, et al.
Published: (2025)
Token-Space Mask Prediction for Efficient Vision Transformer Segmentation
by: Galagain, Calvin, et al.
Published: (2026)
by: Galagain, Calvin, et al.
Published: (2026)
PMT: Progressive Mean Teacher via Exploring Temporal Consistency for Semi-Supervised Medical Image Segmentation
by: Gao, Ning, et al.
Published: (2024)
by: Gao, Ning, et al.
Published: (2024)
Harnessing Frozen Unimodal Encoders for Flexible Multimodal Alignment
by: Maniparambil, Mayug, et al.
Published: (2024)
by: Maniparambil, Mayug, et al.
Published: (2024)
WeakTr: Exploring Plain Vision Transformer for Weakly-supervised Semantic Segmentation
by: Zhu, Lianghui, et al.
Published: (2023)
by: Zhu, Lianghui, et al.
Published: (2023)
CDPDNet: Integrating Text Guidance with Hybrid Vision Encoders for Medical Image Segmentation
by: Wu, Jiong, et al.
Published: (2025)
by: Wu, Jiong, et al.
Published: (2025)
Prompt Generation Networks for Input-Space Adaptation of Frozen Vision Transformers
by: Loedeman, Jochem, et al.
Published: (2022)
by: Loedeman, Jochem, et al.
Published: (2022)
FROST-Drive: Scalable and Efficient End-to-End Driving with a Frozen Vision Encoder
by: Dong, Zeyu, et al.
Published: (2026)
by: Dong, Zeyu, et al.
Published: (2026)
The BRAVO Semantic Segmentation Challenge Results in UNCV2024
by: Vu, Tuan-Hung, et al.
Published: (2024)
by: Vu, Tuan-Hung, et al.
Published: (2024)
Cross-aware Early Fusion with Stage-divided Vision and Language Transformer Encoders for Referring Image Segmentation
by: Cho, Yubin, et al.
Published: (2024)
by: Cho, Yubin, et al.
Published: (2024)
FrozenSeg: Harmonizing Frozen Foundation Models for Open-Vocabulary Segmentation
by: Chen, Xi, et al.
Published: (2024)
by: Chen, Xi, et al.
Published: (2024)
Image-Specific Adaptation of Transformer Encoders for Compute-Efficient Segmentation
by: Yao, Manyi, et al.
Published: (2024)
by: Yao, Manyi, et al.
Published: (2024)
Frozen Transformers in Language Models Are Effective Visual Encoder Layers
by: Pang, Ziqi, et al.
Published: (2023)
by: Pang, Ziqi, et al.
Published: (2023)
Mask4Former: Mask Transformer for 4D Panoptic Segmentation
by: Yilmaz, Kadir, et al.
Published: (2023)
by: Yilmaz, Kadir, et al.
Published: (2023)
Deforming Videos to Masks: Flow Matching for Referring Video Segmentation
by: Wang, Zanyi, et al.
Published: (2025)
by: Wang, Zanyi, et al.
Published: (2025)
Similar Items
-
ALGM: Adaptive Local-then-Global Token Merging for Efficient Semantic Segmentation with Plain Vision Transformers
by: Norouzi, Narges, et al.
Published: (2024) -
VidEoMT: Your ViT is Secretly Also a Video Segmentation Model
by: Norouzi, Narges, et al.
Published: (2026) -
Your ViT is Secretly an Image Segmentation Model
by: Kerssies, Tommie, et al.
Published: (2025) -
Towards Data-Efficient Video Pre-training with Frozen Image Foundation Models
by: Orlova, Svetlana, et al.
Published: (2026) -
Task-aligned Part-aware Panoptic Segmentation through Joint Object-Part Representations
by: de Geus, Daan, et al.
Published: (2024)