Token-Space Mask Prediction for Efficient Vision Transformer Segmentation
Fuente:
arXiv
Saved in:
| Main Authors: | Galagain, Calvin, Poreba, Martyna, Goulette, François |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LiPS: Lightweight Panoptic Segmentation for Resource-Constrained Robotics
by: Galagain, Calvin, et al.
Published: (2026)
by: Galagain, Calvin, et al.
Published: (2026)
Is Semantic SLAM Ready for Embedded Systems ? A Comparative Survey
by: Galagain, Calvin, et al.
Published: (2025)
by: Galagain, Calvin, et al.
Published: (2025)
I-Segmenter: Integer-Only Vision Transformer for Efficient Semantic Segmentation
by: Sassoon, Jordan, et al.
Published: (2025)
by: Sassoon, Jordan, et al.
Published: (2025)
Beyond Attention Scores: SVD-Based Vision Token Pruning for Efficient Vision-Language Models
by: Apedo, Yvon, et al.
Published: (2026)
by: Apedo, Yvon, et al.
Published: (2026)
Where Do Tokens Go? Understanding Pruning Behaviors in STEP at High Resolutions
by: Szczepanski, Michal, et al.
Published: (2025)
by: Szczepanski, Michal, et al.
Published: (2025)
3DLabelProp: Geometric-Driven Domain Generalization for LiDAR Semantic Segmentation in Autonomous Driving
by: Sanchez, Jules, et al.
Published: (2025)
by: Sanchez, Jules, et al.
Published: (2025)
COLA: COarse-LAbel multi-source LiDAR semantic segmentation for autonomous driving
by: Sanchez, Jules, et al.
Published: (2023)
by: Sanchez, Jules, et al.
Published: (2023)
HD-OOD3D: Supervised and Unsupervised Out-of-Distribution object detection in LiDAR data
by: Soum-Fontez, Louis, et al.
Published: (2024)
by: Soum-Fontez, Louis, et al.
Published: (2024)
Medical Referring Image Segmentation via Next-Token Mask Prediction
by: Chen, Xinyu, et al.
Published: (2025)
by: Chen, Xinyu, et al.
Published: (2025)
TCSAFormer: Efficient Vision Transformer with Token Compression and Sparse Attention for Medical Image Segmentation
by: Xia, Zunhui, et al.
Published: (2025)
by: Xia, Zunhui, et al.
Published: (2025)
ALGM: Adaptive Local-then-Global Token Merging for Efficient Semantic Segmentation with Plain Vision Transformers
by: Norouzi, Narges, et al.
Published: (2024)
by: Norouzi, Narges, et al.
Published: (2024)
PPT: Token Pruning and Pooling for Efficient Vision Transformers
by: Wu, Xinjian, et al.
Published: (2023)
by: Wu, Xinjian, et al.
Published: (2023)
PMT: Plain Mask Transformer for Image and Video Segmentation with Frozen Vision Encoders
by: Cavagnero, Niccolò, et al.
Published: (2026)
by: Cavagnero, Niccolò, et al.
Published: (2026)
ToSA: Token Selective Attention for Efficient Vision Transformers
by: Singh, Manish Kumar, et al.
Published: (2024)
by: Singh, Manish Kumar, et al.
Published: (2024)
Adapting Vision Transformers to Ultra-High Resolution Semantic Segmentation with Relay Tokens
by: Perron, Yohann, et al.
Published: (2026)
by: Perron, Yohann, et al.
Published: (2026)
Exploring Token-Level Augmentation in Vision Transformer for Semi-Supervised Semantic Segmentation
by: Zhang, Dengke, et al.
Published: (2025)
by: Zhang, Dengke, et al.
Published: (2025)
Progressive Token Length Scaling in Transformer Encoders for Efficient Universal Segmentation
by: Aich, Abhishek, et al.
Published: (2024)
by: Aich, Abhishek, et al.
Published: (2024)
Towards Efficient Vision State Space Models via Token Merging
by: Park, Jinyoung, et al.
Published: (2025)
by: Park, Jinyoung, et al.
Published: (2025)
ParisLuco3D: A high-quality target dataset for domain generalization of LiDAR perception
by: Sanchez, Jules, et al.
Published: (2023)
by: Sanchez, Jules, et al.
Published: (2023)
Computational Tradeoffs in Image Synthesis: Diffusion, Masked-Token, and Next-Token Prediction
by: Kilian, Maciej, et al.
Published: (2024)
by: Kilian, Maciej, et al.
Published: (2024)
Vision Transformers: From Semantic Segmentation to Dense Prediction
by: Zhang, Li, et al.
Published: (2022)
by: Zhang, Li, et al.
Published: (2022)
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers
by: Bae, Jongseong, et al.
Published: (2024)
by: Bae, Jongseong, et al.
Published: (2024)
Vote&Mix: Plug-and-Play Token Reduction for Efficient Vision Transformer
by: Peng, Shuai, et al.
Published: (2024)
by: Peng, Shuai, et al.
Published: (2024)
Learning to Mask and Permute Visual Tokens for Vision Transformer Pre-Training
by: Baraldi, Lorenzo, et al.
Published: (2023)
by: Baraldi, Lorenzo, et al.
Published: (2023)
Mask4Former: Mask Transformer for 4D Panoptic Segmentation
by: Yilmaz, Kadir, et al.
Published: (2023)
by: Yilmaz, Kadir, et al.
Published: (2023)
Efficient Token Compression for Vision Transformer with Spatial Information Preserved
by: Mao, Junzhu, et al.
Published: (2025)
by: Mao, Junzhu, et al.
Published: (2025)
MaskMed: Decoupled Mask and Class Prediction for Medical Image Segmentation
by: Xie, Bin, et al.
Published: (2025)
by: Xie, Bin, et al.
Published: (2025)
Vision Transformer with Super Token Sampling
by: Huang, Huaibo, et al.
Published: (2022)
by: Huang, Huaibo, et al.
Published: (2022)
EAST: Early Action Prediction Sampling Strategy with Token Masking
by: Sović, Iva, et al.
Published: (2026)
by: Sović, Iva, et al.
Published: (2026)
Efficient Masked Attention Transformer for Few-Shot Classification and Segmentation
by: Carrión-Ojeda, Dustin, et al.
Published: (2025)
by: Carrión-Ojeda, Dustin, et al.
Published: (2025)
Polyline Path Masked Attention for Vision Transformer
by: Zhao, Zhongchen, et al.
Published: (2025)
by: Zhao, Zhongchen, et al.
Published: (2025)
Frequency-Aware Token Reduction for Efficient Vision Transformer
by: Lee, Dong-Jae, et al.
Published: (2025)
by: Lee, Dong-Jae, et al.
Published: (2025)
The Role of Masking for Efficient Supervised Knowledge Distillation of Vision Transformers
by: Son, Seungwoo, et al.
Published: (2023)
by: Son, Seungwoo, et al.
Published: (2023)
GMT: Guided Mask Transformer for Leaf Instance Segmentation
by: Chen, Feng, et al.
Published: (2024)
by: Chen, Feng, et al.
Published: (2024)
MATIS: Masked-Attention Transformers for Surgical Instrument Segmentation
by: Ayobi, Nicolás, et al.
Published: (2023)
by: Ayobi, Nicolás, et al.
Published: (2023)
Vision SmolMamba: Spike-Guided Token Pruning for Energy-Efficient Spiking State-Space Vision Models
by: Bai, Dewei, et al.
Published: (2026)
by: Bai, Dewei, et al.
Published: (2026)
Visual-Word Tokenizer: Beyond Fixed Sets of Tokens in Vision Transformers
by: Gee, Leonidas, et al.
Published: (2024)
by: Gee, Leonidas, et al.
Published: (2024)
Superpixel Tokenization for Vision Transformers: Preserving Semantic Integrity in Visual Tokens
by: Lew, Jaihyun, et al.
Published: (2024)
by: Lew, Jaihyun, et al.
Published: (2024)
HEART-VIT: Hessian-Guided Efficient Dynamic Attention and Token Pruning in Vision Transformer
by: Uddin, Mohammad Helal, et al.
Published: (2025)
by: Uddin, Mohammad Helal, et al.
Published: (2025)
Multi-criteria Token Fusion with One-step-ahead Attention for Efficient Vision Transformers
by: Lee, Sanghyeok, et al.
Published: (2024)
by: Lee, Sanghyeok, et al.
Published: (2024)
Similar Items
-
LiPS: Lightweight Panoptic Segmentation for Resource-Constrained Robotics
by: Galagain, Calvin, et al.
Published: (2026) -
Is Semantic SLAM Ready for Embedded Systems ? A Comparative Survey
by: Galagain, Calvin, et al.
Published: (2025) -
I-Segmenter: Integer-Only Vision Transformer for Efficient Semantic Segmentation
by: Sassoon, Jordan, et al.
Published: (2025) -
Beyond Attention Scores: SVD-Based Vision Token Pruning for Efficient Vision-Language Models
by: Apedo, Yvon, et al.
Published: (2026) -
Where Do Tokens Go? Understanding Pruning Behaviors in STEP at High Resolutions
by: Szczepanski, Michal, et al.
Published: (2025)