SATA: Spatial Autocorrelation Token Analysis for Enhancing the Robustness of Vision Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nikzad, Nick, Liao, Yi, Gao, Yongsheng, Zhou, Jun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CSA-Net: Channel-wise Spatially Autocorrelated Attention Networks
von: Nikzad, Nick, et al.
Veröffentlicht: (2024)
von: Nikzad, Nick, et al.
Veröffentlicht: (2024)
Pruning One More Token is Enough: Leveraging Latency-Workload Non-Linearities for Vision Transformers on the Edge
von: Eliopoulos, Nick John, et al.
Veröffentlicht: (2024)
von: Eliopoulos, Nick John, et al.
Veröffentlicht: (2024)
TraNCE: Transformative Non-linear Concept Explainer for CNNs
von: Akpudo, Ugochukwu Ejike, et al.
Veröffentlicht: (2025)
von: Akpudo, Ugochukwu Ejike, et al.
Veröffentlicht: (2025)
SPoT: Subpixel Placement of Tokens in Vision Transformers
von: Hjelkrem-Tan, Martine, et al.
Veröffentlicht: (2025)
von: Hjelkrem-Tan, Martine, et al.
Veröffentlicht: (2025)
Positional Encodings Anchor Spatial Structure in Vision Transformers: A Geometric Perspective on Robustness
von: Mannes, Mahmoud
Veröffentlicht: (2026)
von: Mannes, Mahmoud
Veröffentlicht: (2026)
Neuron Abandoning Attention Flow: Visual Explanation of Dynamics inside CNN Models
von: Liao, Yi, et al.
Veröffentlicht: (2024)
von: Liao, Yi, et al.
Veröffentlicht: (2024)
DPL: Decoupled Prototype Learning for Enhancing Robustness of Vision-Language Transformers to Missing Modalities
von: Lu, Jueqing, et al.
Veröffentlicht: (2025)
von: Lu, Jueqing, et al.
Veröffentlicht: (2025)
Token Turing Machines are Efficient Vision Models
von: Jajal, Purvish, et al.
Veröffentlicht: (2024)
von: Jajal, Purvish, et al.
Veröffentlicht: (2024)
ZACH-ViT: A Zero-Token Vision Transformer with ShuffleStrides Data Augmentation for Robust Lung Ultrasound Classification
von: Angelakis, Athanasios, et al.
Veröffentlicht: (2025)
von: Angelakis, Athanasios, et al.
Veröffentlicht: (2025)
AdaPerceiver: Transformers with Adaptive Width, Depth, and Tokens
von: Jajal, Purvish, et al.
Veröffentlicht: (2025)
von: Jajal, Purvish, et al.
Veröffentlicht: (2025)
Uncertainty-DTW for Sequences and Visual Tokens
von: Wang, Lei, et al.
Veröffentlicht: (2026)
von: Wang, Lei, et al.
Veröffentlicht: (2026)
Energy-Regularized Spatial Masking: A Novel Approach to Enhancing Robustness and Interpretability in Vision Models
von: Devynck, Tom, et al.
Veröffentlicht: (2026)
von: Devynck, Tom, et al.
Veröffentlicht: (2026)
Robustness Tokens: Towards Adversarial Robustness of Transformers
von: Pulfer, Brian, et al.
Veröffentlicht: (2025)
von: Pulfer, Brian, et al.
Veröffentlicht: (2025)
Leveraging Registers in Vision Transformers for Robust Adaptation
von: Yellapragada, Srikar, et al.
Veröffentlicht: (2025)
von: Yellapragada, Srikar, et al.
Veröffentlicht: (2025)
ToFe: Lagged Token Freezing and Reusing for Efficient Vision Transformer Inference
von: Zhang, Haoyue, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyue, et al.
Veröffentlicht: (2025)
Inducing Spatial Locality in Vision Transformers through the Training Protocol
von: Toledo, Eduardo Santiago, et al.
Veröffentlicht: (2026)
von: Toledo, Eduardo Santiago, et al.
Veröffentlicht: (2026)
From Edges to Depth: Probing the Spatial Hierarchy in Vision Transformers
von: Sanghavi, Jainum
Veröffentlicht: (2026)
von: Sanghavi, Jainum
Veröffentlicht: (2026)
Approximate Nullspace Augmented Finetuning for Robust Vision Transformers
von: Liu, Haoyang, et al.
Veröffentlicht: (2024)
von: Liu, Haoyang, et al.
Veröffentlicht: (2024)
Dynamic Accumulated Attention Map for Interpreting Evolution of Decision-Making in Vision Transformer
von: Liao, Yi, et al.
Veröffentlicht: (2025)
von: Liao, Yi, et al.
Veröffentlicht: (2025)
Mechanistic Understandings of Representation Vulnerabilities and Engineering Robust Vision Transformers
von: Islam, Chashi Mahiul, et al.
Veröffentlicht: (2025)
von: Islam, Chashi Mahiul, et al.
Veröffentlicht: (2025)
Enhancing JEPAs with Spatial Conditioning: Robust and Efficient Representation Learning
von: Littwin, Etai, et al.
Veröffentlicht: (2024)
von: Littwin, Etai, et al.
Veröffentlicht: (2024)
MABViT -- Modified Attention Block Enhances Vision Transformers
von: Ramesh, Mahesh, et al.
Veröffentlicht: (2023)
von: Ramesh, Mahesh, et al.
Veröffentlicht: (2023)
CMAViT: Integrating Climate, Managment, and Remote Sensing Data for Crop Yield Estimation with Multimodel Vision Transformers
von: Kamangir, Hamid, et al.
Veröffentlicht: (2024)
von: Kamangir, Hamid, et al.
Veröffentlicht: (2024)
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training
von: Chen, Yangyi, et al.
Veröffentlicht: (2025)
von: Chen, Yangyi, et al.
Veröffentlicht: (2025)
Token Caching for Diffusion Transformer Acceleration
von: Lou, Jinming, et al.
Veröffentlicht: (2024)
von: Lou, Jinming, et al.
Veröffentlicht: (2024)
LayerShuffle: Enhancing Robustness in Vision Transformers by Randomizing Layer Execution Order
von: Freiberger, Matthias, et al.
Veröffentlicht: (2024)
von: Freiberger, Matthias, et al.
Veröffentlicht: (2024)
Privacy-Aware Video Anomaly Detection through Orthogonal Subspace Projection
von: Wang, Lei, et al.
Veröffentlicht: (2026)
von: Wang, Lei, et al.
Veröffentlicht: (2026)
ELSA: Exact Linear-Scan Attention for Fast and Memory-Light Vision Transformers
von: Hsu, Chih-Chung, et al.
Veröffentlicht: (2026)
von: Hsu, Chih-Chung, et al.
Veröffentlicht: (2026)
Split Adaptation for Pre-trained Vision Transformers
von: Wang, Lixu, et al.
Veröffentlicht: (2025)
von: Wang, Lixu, et al.
Veröffentlicht: (2025)
Efficient Visual Transformer by Learnable Token Merging
von: Wang, Yancheng, et al.
Veröffentlicht: (2024)
von: Wang, Yancheng, et al.
Veröffentlicht: (2024)
Trust-Aware Joint Feature-Prediction Discrepancy for Robust Domain Adaptation
von: Ding, Xi, et al.
Veröffentlicht: (2026)
von: Ding, Xi, et al.
Veröffentlicht: (2026)
SpecFormer: Guarding Vision Transformer Robustness via Maximum Singular Value Penalization
von: Hu, Xixu, et al.
Veröffentlicht: (2024)
von: Hu, Xixu, et al.
Veröffentlicht: (2024)
TORE: Token Recycling in Vision Transformers for Efficient Active Visual Exploration
von: Olszewski, Jan, et al.
Veröffentlicht: (2023)
von: Olszewski, Jan, et al.
Veröffentlicht: (2023)
Data-independent Module-aware Pruning for Hierarchical Vision Transformers
von: He, Yang, et al.
Veröffentlicht: (2024)
von: He, Yang, et al.
Veröffentlicht: (2024)
HQViT: Hybrid Quantum Vision Transformer for Image Classification
von: Zhang, Hui, et al.
Veröffentlicht: (2025)
von: Zhang, Hui, et al.
Veröffentlicht: (2025)
Unlocking Noise-Resistant Vision: Key Architectural Secrets for Robust Models
von: Kim, Bum Jun, et al.
Veröffentlicht: (2025)
von: Kim, Bum Jun, et al.
Veröffentlicht: (2025)
Integrating ConvNeXt and Vision Transformers for Enhancing Facial Age Estimation
von: Maroun, Gaby, et al.
Veröffentlicht: (2025)
von: Maroun, Gaby, et al.
Veröffentlicht: (2025)
Propensity-driven Uncertainty Learning for Sample Exploration in Source-Free Active Domain Adaptation
von: Pan, Zicheng, et al.
Veröffentlicht: (2025)
von: Pan, Zicheng, et al.
Veröffentlicht: (2025)
A Comparative Survey of Vision Transformers for Feature Extraction in Texture Analysis
von: Scabini, Leonardo, et al.
Veröffentlicht: (2024)
von: Scabini, Leonardo, et al.
Veröffentlicht: (2024)
CubistMerge: Spatial-Preserving Token Merging For Diverse ViT Backbones
von: Gong, Wenyi, et al.
Veröffentlicht: (2025)
von: Gong, Wenyi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CSA-Net: Channel-wise Spatially Autocorrelated Attention Networks
von: Nikzad, Nick, et al.
Veröffentlicht: (2024) -
Pruning One More Token is Enough: Leveraging Latency-Workload Non-Linearities for Vision Transformers on the Edge
von: Eliopoulos, Nick John, et al.
Veröffentlicht: (2024) -
TraNCE: Transformative Non-linear Concept Explainer for CNNs
von: Akpudo, Ugochukwu Ejike, et al.
Veröffentlicht: (2025) -
SPoT: Subpixel Placement of Tokens in Vision Transformers
von: Hjelkrem-Tan, Martine, et al.
Veröffentlicht: (2025) -
Positional Encodings Anchor Spatial Structure in Vision Transformers: A Geometric Perspective on Robustness
von: Mannes, Mahmoud
Veröffentlicht: (2026)