Revitalizing Dense Material Segmentation: Stabilized Vision Transformers and the Generalization Paradox
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kazakov, Allan, Cakir, Duygu, İrfanoğlu, Hilal Kurt, İrfanoğlu, Yavuz |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Vision Transformers: From Semantic Segmentation to Dense Prediction
von: Zhang, Li, et al.
Veröffentlicht: (2022)
von: Zhang, Li, et al.
Veröffentlicht: (2022)
Quantization Robustness to Input Degradations for Object Detection
von: Karimov, Toghrul, et al.
Veröffentlicht: (2025)
von: Karimov, Toghrul, et al.
Veröffentlicht: (2025)
Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis
von: Bai, Jinbin, et al.
Veröffentlicht: (2024)
von: Bai, Jinbin, et al.
Veröffentlicht: (2024)
Dense Vision Transformer Compression with Few Samples
von: Zhang, Hanxiao, et al.
Veröffentlicht: (2024)
von: Zhang, Hanxiao, et al.
Veröffentlicht: (2024)
RD-ViT: Recurrent-Depth Vision Transformer for Semantic Segmentation with Reduced Data Dependence Extending the Recurrent-Depth Transformer Architecture to Dense Prediction
von: He, Renjie
Veröffentlicht: (2026)
von: He, Renjie
Veröffentlicht: (2026)
VPNeXt -- Rethinking Dense Decoding for Plain Vision Transformer
von: Tang, Xikai, et al.
Veröffentlicht: (2025)
von: Tang, Xikai, et al.
Veröffentlicht: (2025)
CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction
von: Wu, Size, et al.
Veröffentlicht: (2023)
von: Wu, Size, et al.
Veröffentlicht: (2023)
MaterialSeg3D: Segmenting Dense Materials from 2D Priors for 3D Assets
von: Li, Zeyu, et al.
Veröffentlicht: (2024)
von: Li, Zeyu, et al.
Veröffentlicht: (2024)
The Power of Certainty: How Confident Models Lead to Better Segmentation
von: Erol, Tugberk, et al.
Veröffentlicht: (2025)
von: Erol, Tugberk, et al.
Veröffentlicht: (2025)
SHED Light on Segmentation for Dense Prediction
von: Lee, Seung Hyun, et al.
Veröffentlicht: (2026)
von: Lee, Seung Hyun, et al.
Veröffentlicht: (2026)
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer
von: Huang, Shaofei, et al.
Veröffentlicht: (2025)
von: Huang, Shaofei, et al.
Veröffentlicht: (2025)
ViT-CoMer: Vision Transformer with Convolutional Multi-scale Feature Interaction for Dense Predictions
von: Xia, Chunlong, et al.
Veröffentlicht: (2024)
von: Xia, Chunlong, et al.
Veröffentlicht: (2024)
DenSe-AdViT: A novel Vision Transformer for Dense SAR Object Detection
von: Zhang, Yang, et al.
Veröffentlicht: (2025)
von: Zhang, Yang, et al.
Veröffentlicht: (2025)
Native Segmentation Vision Transformers
von: Brasó, Guillem, et al.
Veröffentlicht: (2025)
von: Brasó, Guillem, et al.
Veröffentlicht: (2025)
DSGG: Dense Relation Transformer for an End-to-end Scene Graph Generation
von: Hayder, Zeeshan, et al.
Veröffentlicht: (2024)
von: Hayder, Zeeshan, et al.
Veröffentlicht: (2024)
Weakly Supervised Food Image Segmentation using Vision Transformers and Segment Anything Model
von: Sarafis, Ioannis, et al.
Veröffentlicht: (2025)
von: Sarafis, Ioannis, et al.
Veröffentlicht: (2025)
Dino-NestedUNet: Unlocking Foundation Vision Encoders for Pathology Tumor Bulk Segmentation via Dense Decoding
von: Wang, Tianyang, et al.
Veröffentlicht: (2026)
von: Wang, Tianyang, et al.
Veröffentlicht: (2026)
Token-Space Mask Prediction for Efficient Vision Transformer Segmentation
von: Galagain, Calvin, et al.
Veröffentlicht: (2026)
von: Galagain, Calvin, et al.
Veröffentlicht: (2026)
Boundary-Aware Vision Transformer for Angiography Vascular Network Segmentation
von: Hezil, Nabil, et al.
Veröffentlicht: (2025)
von: Hezil, Nabil, et al.
Veröffentlicht: (2025)
SAMIDARE: Advanced Tracking-by-Segmentation for Dense Scenarios
von: Hirano, Shozaburo, et al.
Veröffentlicht: (2026)
von: Hirano, Shozaburo, et al.
Veröffentlicht: (2026)
Virtual Category Learning: A Semi-Supervised Learning Method for Dense Prediction with Extremely Limited Labels
von: Chen, Changrui, et al.
Veröffentlicht: (2023)
von: Chen, Changrui, et al.
Veröffentlicht: (2023)
DisentangleFormer: Spatial-Channel Decoupling for Multi-Channel Vision
von: Liao, Jiashu, et al.
Veröffentlicht: (2025)
von: Liao, Jiashu, et al.
Veröffentlicht: (2025)
Diff-Plugin: Revitalizing Details for Diffusion-based Low-level Tasks
von: Liu, Yuhao, et al.
Veröffentlicht: (2024)
von: Liu, Yuhao, et al.
Veröffentlicht: (2024)
DiCo: Revitalizing ConvNets for Scalable and Efficient Diffusion Modeling
von: Ai, Yuang, et al.
Veröffentlicht: (2025)
von: Ai, Yuang, et al.
Veröffentlicht: (2025)
Unified Review and Benchmark of Deep Segmentation Architectures for Cardiac Ultrasound on CAMUS
von: Ullah, Zahid, et al.
Veröffentlicht: (2025)
von: Ullah, Zahid, et al.
Veröffentlicht: (2025)
Neuromorphic Vision-based Motion Segmentation with Graph Transformer Neural Network
von: Alkendi, Yusra, et al.
Veröffentlicht: (2024)
von: Alkendi, Yusra, et al.
Veröffentlicht: (2024)
Self-Supervised Vision Transformers Are Efficient Segmentation Learners for Imperfect Labels
von: Lee, Seungho, et al.
Veröffentlicht: (2024)
von: Lee, Seungho, et al.
Veröffentlicht: (2024)
AViT: Adapting Vision Transformers for Small Skin Lesion Segmentation Datasets
von: Du, Siyi, et al.
Veröffentlicht: (2023)
von: Du, Siyi, et al.
Veröffentlicht: (2023)
Multi-Layer Dense Attention Decoder for Polyp Segmentation
von: Patel, Krushi, et al.
Veröffentlicht: (2024)
von: Patel, Krushi, et al.
Veröffentlicht: (2024)
Mitigating Data Redundancy to Revitalize Transformer-based Long-Term Time Series Forecasting System
von: Li, Mingjie, et al.
Veröffentlicht: (2022)
von: Li, Mingjie, et al.
Veröffentlicht: (2022)
DenseSeg: Joint Learning for Semantic Segmentation and Landmark Detection Using Dense Image-to-Shape Representation
von: Keuth, Ron, et al.
Veröffentlicht: (2024)
von: Keuth, Ron, et al.
Veröffentlicht: (2024)
DBAT: Dynamic Backward Attention Transformer for Material Segmentation with Cross-Resolution Patches
von: Heng, Yuwen, et al.
Veröffentlicht: (2023)
von: Heng, Yuwen, et al.
Veröffentlicht: (2023)
Memory Efficient Transformer Adapter for Dense Predictions
von: Zhang, Dong, et al.
Veröffentlicht: (2025)
von: Zhang, Dong, et al.
Veröffentlicht: (2025)
MMSFormer: Multimodal Transformer for Material and Semantic Segmentation
von: Reza, Md Kaykobad, et al.
Veröffentlicht: (2023)
von: Reza, Md Kaykobad, et al.
Veröffentlicht: (2023)
CroMo-Mixup: Augmenting Cross-Model Representations for Continual Self-Supervised Learning
von: Mushtaq, Erum, et al.
Veröffentlicht: (2024)
von: Mushtaq, Erum, et al.
Veröffentlicht: (2024)
PlankFormer: Robust Plankton Instance Segmentation via MAE-Pretrained Vision Transformers and Pseudo Community Image Generation
von: Miyazaki, Masaharu, et al.
Veröffentlicht: (2026)
von: Miyazaki, Masaharu, et al.
Veröffentlicht: (2026)
Forest2Seq: Revitalizing Order Prior for Sequential Indoor Scene Synthesis
von: Sun, Qi, et al.
Veröffentlicht: (2024)
von: Sun, Qi, et al.
Veröffentlicht: (2024)
Adapting Vision Transformers to Ultra-High Resolution Semantic Segmentation with Relay Tokens
von: Perron, Yohann, et al.
Veröffentlicht: (2026)
von: Perron, Yohann, et al.
Veröffentlicht: (2026)
PMT: Plain Mask Transformer for Image and Video Segmentation with Frozen Vision Encoders
von: Cavagnero, Niccolò, et al.
Veröffentlicht: (2026)
von: Cavagnero, Niccolò, et al.
Veröffentlicht: (2026)
WeakTr: Exploring Plain Vision Transformer for Weakly-supervised Semantic Segmentation
von: Zhu, Lianghui, et al.
Veröffentlicht: (2023)
von: Zhu, Lianghui, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Vision Transformers: From Semantic Segmentation to Dense Prediction
von: Zhang, Li, et al.
Veröffentlicht: (2022) -
Quantization Robustness to Input Degradations for Object Detection
von: Karimov, Toghrul, et al.
Veröffentlicht: (2025) -
Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis
von: Bai, Jinbin, et al.
Veröffentlicht: (2024) -
Dense Vision Transformer Compression with Few Samples
von: Zhang, Hanxiao, et al.
Veröffentlicht: (2024) -
RD-ViT: Recurrent-Depth Vision Transformer for Semantic Segmentation with Reduced Data Dependence Extending the Recurrent-Depth Transformer Architecture to Dense Prediction
von: He, Renjie
Veröffentlicht: (2026)