Saved in:
| Main Author: | Sanghavi, Jainum |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.23452 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Deep Learning for Melt Pool Depth Contour Prediction From Surface Thermal Images via Vision Transformers
by: Ogoke, Francis, et al.
Published: (2024)
by: Ogoke, Francis, et al.
Published: (2024)
Inducing Spatial Locality in Vision Transformers through the Training Protocol
by: Toledo, Eduardo Santiago, et al.
Published: (2026)
by: Toledo, Eduardo Santiago, et al.
Published: (2026)
Light Cones For Vision: Simple Causal Priors For Visual Hierarchy
by: Kartik, Manglam, et al.
Published: (2026)
by: Kartik, Manglam, et al.
Published: (2026)
SATA: Spatial Autocorrelation Token Analysis for Enhancing the Robustness of Vision Transformers
by: Nikzad, Nick, et al.
Published: (2024)
by: Nikzad, Nick, et al.
Published: (2024)
From Colors to Classes: Emergence of Concepts in Vision Transformers
by: Dorszewski, Teresa, et al.
Published: (2025)
by: Dorszewski, Teresa, et al.
Published: (2025)
Positional Encodings Anchor Spatial Structure in Vision Transformers: A Geometric Perspective on Robustness
by: Mannes, Mahmoud
Published: (2026)
by: Mannes, Mahmoud
Published: (2026)
Leaner Transformers: More Heads, Less Depth
by: Saratchandran, Hemanth, et al.
Published: (2025)
by: Saratchandran, Hemanth, et al.
Published: (2025)
AdaPerceiver: Transformers with Adaptive Width, Depth, and Tokens
by: Jajal, Purvish, et al.
Published: (2025)
by: Jajal, Purvish, et al.
Published: (2025)
PHyCLIP: $\ell_1$-Product of Hyperbolic Factors Unifies Hierarchy and Compositionality in Vision-Language Representation Learning
by: Yoshikawa, Daiki, et al.
Published: (2025)
by: Yoshikawa, Daiki, et al.
Published: (2025)
Probing the Representational Power of Sparse Autoencoders in Vision Models
by: Olson, Matthew Lyle, et al.
Published: (2025)
by: Olson, Matthew Lyle, et al.
Published: (2025)
From Logits to Hierarchies: Hierarchical Clustering made Simple
by: Palumbo, Emanuele, et al.
Published: (2024)
by: Palumbo, Emanuele, et al.
Published: (2024)
Native Segmentation Vision Transformers
by: Brasó, Guillem, et al.
Published: (2025)
by: Brasó, Guillem, et al.
Published: (2025)
Advancing Autonomous Driving: DepthSense with Radar and Spatial Attention
by: Hussain, Muhamamd Ishfaq, et al.
Published: (2021)
by: Hussain, Muhamamd Ishfaq, et al.
Published: (2021)
The Dual Mechanisms of Spatial Reasoning in Vision-Language Models
by: Cui, Kelly, et al.
Published: (2026)
by: Cui, Kelly, et al.
Published: (2026)
Iwin Transformer: Hierarchical Vision Transformer using Interleaved Windows
by: Huo, Simin, et al.
Published: (2025)
by: Huo, Simin, et al.
Published: (2025)
RAViT: Resolution-Adaptive Vision Transformer
by: Guidez, Martial, et al.
Published: (2026)
by: Guidez, Martial, et al.
Published: (2026)
Slicing Vision Transformer for Flexible Inference
by: Zhang, Yitian, et al.
Published: (2024)
by: Zhang, Yitian, et al.
Published: (2024)
Rotary Position Embedding for Vision Transformer
by: Heo, Byeongho, et al.
Published: (2024)
by: Heo, Byeongho, et al.
Published: (2024)
Decorrelation Speeds Up Vision Transformers
by: Carrigg, Kieran, et al.
Published: (2025)
by: Carrigg, Kieran, et al.
Published: (2025)
DRWKV: Focusing on Object Edges for Low-Light Image Enhancement
by: Bai, Xuecheng, et al.
Published: (2025)
by: Bai, Xuecheng, et al.
Published: (2025)
HaloProbe: Bayesian Detection and Mitigation of Object Hallucinations in Vision-Language Models
by: Zohrabi, Reihaneh, et al.
Published: (2026)
by: Zohrabi, Reihaneh, et al.
Published: (2026)
Hierarchy-Guided Multimodal Representation Learning for Taxonomic Inference
by: Ahmed, Sk Miraj, et al.
Published: (2026)
by: Ahmed, Sk Miraj, et al.
Published: (2026)
From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
by: Zhang, Zhengshen, et al.
Published: (2025)
by: Zhang, Zhengshen, et al.
Published: (2025)
Elastic Attention Cores for Scalable Vision Transformers
by: Song, Alan Z., et al.
Published: (2026)
by: Song, Alan Z., et al.
Published: (2026)
Attention Transfer Is Not Universally Effective for Vision Transformers
by: Qin, Huaiyuan, et al.
Published: (2026)
by: Qin, Huaiyuan, et al.
Published: (2026)
SPoT: Subpixel Placement of Tokens in Vision Transformers
by: Hjelkrem-Tan, Martine, et al.
Published: (2025)
by: Hjelkrem-Tan, Martine, et al.
Published: (2025)
Instance-Aware Group Quantization for Vision Transformers
by: Moon, Jaehyeon, et al.
Published: (2024)
by: Moon, Jaehyeon, et al.
Published: (2024)
Vision Transformer-based Adversarial Domain Adaptation
by: Li, Yahan, et al.
Published: (2024)
by: Li, Yahan, et al.
Published: (2024)
Compact Vision Transformer by Reduction of Kernel Complexity
by: Wang, Yancheng, et al.
Published: (2025)
by: Wang, Yancheng, et al.
Published: (2025)
Split Adaptation for Pre-trained Vision Transformers
by: Wang, Lixu, et al.
Published: (2025)
by: Wang, Lixu, et al.
Published: (2025)
DASViT: Differentiable Architecture Search for Vision Transformer
by: Wu, Pengjin, et al.
Published: (2025)
by: Wu, Pengjin, et al.
Published: (2025)
Stable Vision Concept Transformers for Medical Diagnosis
by: Hu, Lijie, et al.
Published: (2025)
by: Hu, Lijie, et al.
Published: (2025)
Self-Supervised Vision Transformers for Writer Retrieval
by: Raven, Tim, et al.
Published: (2024)
by: Raven, Tim, et al.
Published: (2024)
HEAL-SWIN: A Vision Transformer On The Sphere
by: Carlsson, Oscar, et al.
Published: (2023)
by: Carlsson, Oscar, et al.
Published: (2023)
A Manifold Representation of the Key in Vision Transformers
by: Meng, Li, et al.
Published: (2024)
by: Meng, Li, et al.
Published: (2024)
Leveraging Registers in Vision Transformers for Robust Adaptation
by: Yellapragada, Srikar, et al.
Published: (2025)
by: Yellapragada, Srikar, et al.
Published: (2025)
Geometric Flood Depth Estimation: Fusing Transformer-Based Segmentation with Digital Elevation Models
by: Le, Nhut, et al.
Published: (2026)
by: Le, Nhut, et al.
Published: (2026)
MentalBlackboard: Evaluating Spatial Visualization via Mathematical Transformations
by: Yilmaz, Nilay, et al.
Published: (2026)
by: Yilmaz, Nilay, et al.
Published: (2026)
Head Pursuit: Probing Attention Specialization in Multimodal Transformers
by: Basile, Lorenzo, et al.
Published: (2025)
by: Basile, Lorenzo, et al.
Published: (2025)
Do VLMs Need Vision Transformers? Evaluating State Space Models as Vision Encoders
by: Kuo, Shang-Jui Ray, et al.
Published: (2026)
by: Kuo, Shang-Jui Ray, et al.
Published: (2026)
Similar Items
-
Deep Learning for Melt Pool Depth Contour Prediction From Surface Thermal Images via Vision Transformers
by: Ogoke, Francis, et al.
Published: (2024) -
Inducing Spatial Locality in Vision Transformers through the Training Protocol
by: Toledo, Eduardo Santiago, et al.
Published: (2026) -
Light Cones For Vision: Simple Causal Priors For Visual Hierarchy
by: Kartik, Manglam, et al.
Published: (2026) -
SATA: Spatial Autocorrelation Token Analysis for Enhancing the Robustness of Vision Transformers
by: Nikzad, Nick, et al.
Published: (2024) -
From Colors to Classes: Emergence of Concepts in Vision Transformers
by: Dorszewski, Teresa, et al.
Published: (2025)