Adapting Vision Transformers to Ultra-High Resolution Semantic Segmentation with Relay Tokens
Fuente:
arXiv
Saved in:
| Main Authors: | Perron, Yohann, Sydorov, Vladyslav, Pottier, Christophe, Landrieu, Loic |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Archaeoscape: Bringing Aerial Laser Scanning Archaeology to the Deep Learning Era
by: Perron, Yohann, et al.
Published: (2024)
by: Perron, Yohann, et al.
Published: (2024)
Scalable 3D Panoptic Segmentation As Superpoint Graph Clustering
by: Robert, Damien, et al.
Published: (2024)
by: Robert, Damien, et al.
Published: (2024)
EZ-SP: Fast and Lightweight Superpoint-Based 3D Segmentation
by: Geist, Louis, et al.
Published: (2025)
by: Geist, Louis, et al.
Published: (2025)
Open-Canopy: Towards Very High Resolution Forest Monitoring
by: Fogel, Fajwel, et al.
Published: (2024)
by: Fogel, Fajwel, et al.
Published: (2024)
AnySat: One Earth Observation Model for Many Resolutions, Scales, and Modalities
by: Astruc, Guillaume, et al.
Published: (2024)
by: Astruc, Guillaume, et al.
Published: (2024)
Segmenting France Across Four Centuries
by: López-Rauhut, Marta, et al.
Published: (2025)
by: López-Rauhut, Marta, et al.
Published: (2025)
Exploring Token-Level Augmentation in Vision Transformer for Semi-Supervised Semantic Segmentation
by: Zhang, Dengke, et al.
Published: (2025)
by: Zhang, Dengke, et al.
Published: (2025)
Towards Robust and Generalizable Lensless Imaging with Modular Learned Reconstruction
by: Bezzam, Eric, et al.
Published: (2025)
by: Bezzam, Eric, et al.
Published: (2025)
Learnable Earth Parser: Discovering 3D Prototypes in Aerial Scans
by: Loiseau, Romain, et al.
Published: (2023)
by: Loiseau, Romain, et al.
Published: (2023)
CoDEx: Combining Domain Expertise for Spatial Generalization in Satellite Image Analysis
by: Kuriyal, Abhishek, et al.
Published: (2025)
by: Kuriyal, Abhishek, et al.
Published: (2025)
OmniSat: Self-Supervised Modality Fusion for Earth Observation
by: Astruc, Guillaume, et al.
Published: (2024)
by: Astruc, Guillaume, et al.
Published: (2024)
A Survey and Benchmark of Automatic Surface Reconstruction from Point Clouds
by: Sulzer, Raphael, et al.
Published: (2023)
by: Sulzer, Raphael, et al.
Published: (2023)
Ultra-High Resolution Segmentation via Boundary-Enhanced Patch-Merging Transformer
by: Sun, Haopeng, et al.
Published: (2024)
by: Sun, Haopeng, et al.
Published: (2024)
ALGM: Adaptive Local-then-Global Token Merging for Efficient Semantic Segmentation with Plain Vision Transformers
by: Norouzi, Narges, et al.
Published: (2024)
by: Norouzi, Narges, et al.
Published: (2024)
Superpixel Tokenization for Vision Transformers: Preserving Semantic Integrity in Visual Tokens
by: Lew, Jaihyun, et al.
Published: (2024)
by: Lew, Jaihyun, et al.
Published: (2024)
Adapting Vision-Language Model with Fine-grained Semantics for Open-Vocabulary Segmentation
by: Chng, Yong Xien, et al.
Published: (2024)
by: Chng, Yong Xien, et al.
Published: (2024)
AViT: Adapting Vision Transformers for Small Skin Lesion Segmentation Datasets
by: Du, Siyi, et al.
Published: (2023)
by: Du, Siyi, et al.
Published: (2023)
Order Matters: 3D Shape Generation from Sequential VR Sketches
by: Chen, Yizi, et al.
Published: (2025)
by: Chen, Yizi, et al.
Published: (2025)
RAViT: Resolution-Adaptive Vision Transformer
by: Guidez, Martial, et al.
Published: (2026)
by: Guidez, Martial, et al.
Published: (2026)
Around the World in 80 Timesteps: A Generative Approach to Global Visual Geolocation
by: Dufour, Nicolas, et al.
Published: (2024)
by: Dufour, Nicolas, et al.
Published: (2024)
Token-Space Mask Prediction for Efficient Vision Transformer Segmentation
by: Galagain, Calvin, et al.
Published: (2026)
by: Galagain, Calvin, et al.
Published: (2026)
Segformer++: Efficient Token-Merging Strategies for High-Resolution Semantic Segmentation
by: Kienzle, Daniel, et al.
Published: (2024)
by: Kienzle, Daniel, et al.
Published: (2024)
Sparse Refinement for Efficient High-Resolution Semantic Segmentation
by: Liu, Zhijian, et al.
Published: (2024)
by: Liu, Zhijian, et al.
Published: (2024)
UNIGEOCLIP: Unified Geospatial Contrastive Learning
by: Astruc, Guillaume, et al.
Published: (2026)
by: Astruc, Guillaume, et al.
Published: (2026)
Vision Transformers: From Semantic Segmentation to Dense Prediction
by: Zhang, Li, et al.
Published: (2022)
by: Zhang, Li, et al.
Published: (2022)
Privacy-Preserving Semantic Segmentation from Ultra-Low-Resolution RGB Inputs
by: Huang, Xuying, et al.
Published: (2025)
by: Huang, Xuying, et al.
Published: (2025)
Adapting LLaMA Decoder to Vision Transformer
by: Wang, Jiahao, et al.
Published: (2024)
by: Wang, Jiahao, et al.
Published: (2024)
RHRSegNet: Relighting High-Resolution Night-Time Semantic Segmentation
by: Elmahdy, Sarah, et al.
Published: (2024)
by: Elmahdy, Sarah, et al.
Published: (2024)
On the Effect of Image Resolution on Semantic Segmentation
by: Singh, Ritambhara, et al.
Published: (2024)
by: Singh, Ritambhara, et al.
Published: (2024)
UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing
by: Dang, Yunkai, et al.
Published: (2026)
by: Dang, Yunkai, et al.
Published: (2026)
TCSAFormer: Efficient Vision Transformer with Token Compression and Sparse Attention for Medical Image Segmentation
by: Xia, Zunhui, et al.
Published: (2025)
by: Xia, Zunhui, et al.
Published: (2025)
Tokenizing Semantic Segmentation with Run Length Encoding
by: Singh, Abhineet, et al.
Published: (2026)
by: Singh, Abhineet, et al.
Published: (2026)
Representation Separation for Semantic Segmentation with Vision Transformers
by: Hong, Yuanduo, et al.
Published: (2022)
by: Hong, Yuanduo, et al.
Published: (2022)
Leveraging Adaptive Implicit Representation Mapping for Ultra High-Resolution Image Segmentation
by: Zhao, Ziyu, et al.
Published: (2024)
by: Zhao, Ziyu, et al.
Published: (2024)
Pyramid Token Pruning for High-Resolution Large Vision-Language Models via Region, Token, and Instruction-Guided Importance
by: Liang, Yuxuan, et al.
Published: (2025)
by: Liang, Yuxuan, et al.
Published: (2025)
Semantic-Geometric Dual Compression: Training-Free Visual Token Reduction for Ultra-High-Resolution Remote Sensing Understanding
by: Li, Yueying, et al.
Published: (2026)
by: Li, Yueying, et al.
Published: (2026)
SAM2-UNeXT: An Improved High-Resolution Baseline for Adapting Foundation Models to Downstream Segmentation Tasks
by: Xiong, Xinyu, et al.
Published: (2025)
by: Xiong, Xinyu, et al.
Published: (2025)
Behind Every Domain There is a Shift: Adapting Distortion-aware Vision Transformers for Panoramic Semantic Segmentation
by: Zhang, Jiaming, et al.
Published: (2022)
by: Zhang, Jiaming, et al.
Published: (2022)
Real-time High-Resolution Neural Network with Semantic Guidance for Crack Segmentation
by: Li, Yongshang, et al.
Published: (2023)
by: Li, Yongshang, et al.
Published: (2023)
Learning to Adapt to Position Bias in Vision Transformer Classifiers
by: Bruintjes, Robert-Jan, et al.
Published: (2025)
by: Bruintjes, Robert-Jan, et al.
Published: (2025)
Similar Items
-
Archaeoscape: Bringing Aerial Laser Scanning Archaeology to the Deep Learning Era
by: Perron, Yohann, et al.
Published: (2024) -
Scalable 3D Panoptic Segmentation As Superpoint Graph Clustering
by: Robert, Damien, et al.
Published: (2024) -
EZ-SP: Fast and Lightweight Superpoint-Based 3D Segmentation
by: Geist, Louis, et al.
Published: (2025) -
Open-Canopy: Towards Very High Resolution Forest Monitoring
by: Fogel, Fajwel, et al.
Published: (2024) -
AnySat: One Earth Observation Model for Many Resolutions, Scales, and Modalities
by: Astruc, Guillaume, et al.
Published: (2024)