Octic Vision Transformers: Quicker ViTs Through Equivariance
Fuente:
arXiv
Saved in:
| Main Authors: | Nordström, David, Edstedt, Johan, Kahl, Fredrik, Bökman, Georg |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MuM: Multi-View Masked Image Modeling for 3D Vision
by: Nordström, David, et al.
Published: (2025)
by: Nordström, David, et al.
Published: (2025)
Who Handles Orientation? Investigating Invariance in Feature Matching
by: Nordström, David, et al.
Published: (2026)
by: Nordström, David, et al.
Published: (2026)
Flopping for FLOPs: Leveraging equivariance for computational efficiency
by: Bökman, Georg, et al.
Published: (2025)
by: Bökman, Georg, et al.
Published: (2025)
TAP-ViTs: Task-Adaptive Pruning for On-Device Deployment of Vision Transformers
by: Wang, Zhibo, et al.
Published: (2026)
by: Wang, Zhibo, et al.
Published: (2026)
Steerers: A framework for rotation equivariant keypoint descriptors
by: Bökman, Georg, et al.
Published: (2023)
by: Bökman, Georg, et al.
Published: (2023)
Affine steerers for structured keypoint description
by: Bökman, Georg, et al.
Published: (2024)
by: Bökman, Georg, et al.
Published: (2024)
Training-Free Acceleration of ViTs with Delayed Spatial Merging
by: Heo, Jung Hwan, et al.
Published: (2023)
by: Heo, Jung Hwan, et al.
Published: (2023)
ConcatPlexer: Additional Dim1 Batching for Faster ViTs
by: Han, Donghoon, et al.
Published: (2023)
by: Han, Donghoon, et al.
Published: (2023)
Communication Efficient Split Learning of ViTs with Attention-based Double Compression
by: Alvetreti, Federico, et al.
Published: (2025)
by: Alvetreti, Federico, et al.
Published: (2025)
Concept-Guided Fine-Tuning: Steering ViTs away from Spurious Correlations to Improve Robustness
by: Elisha, Yehonatan, et al.
Published: (2026)
by: Elisha, Yehonatan, et al.
Published: (2026)
Purrturbed but Stable: Human-Cat Invariant Representations Across CNNs, ViTs and Self-Supervised ViTs
by: Shah, Arya, et al.
Published: (2025)
by: Shah, Arya, et al.
Published: (2025)
STRAP-ViT: Segregated Tokens with Randomized -- Transformations for Defense against Adversarial Patches in ViTs
by: Chattopadhyay, Nandish, et al.
Published: (2026)
by: Chattopadhyay, Nandish, et al.
Published: (2026)
Quasar-ViT: Hardware-Oriented Quantization-Aware Architecture Search for Vision Transformers
by: Li, Zhengang, et al.
Published: (2024)
by: Li, Zhengang, et al.
Published: (2024)
ViTs are Everywhere: A Comprehensive Study Showcasing Vision Transformers in Different Domain
by: Mia, Md Sohag, et al.
Published: (2023)
by: Mia, Md Sohag, et al.
Published: (2023)
HydraViT: Stacking Heads for a Scalable ViT
by: Haberer, Janek, et al.
Published: (2024)
by: Haberer, Janek, et al.
Published: (2024)
TPC-ViT: Token Propagation Controller for Efficient Vision Transformer
by: Zhu, Wentao
Published: (2024)
by: Zhu, Wentao
Published: (2024)
LoMa: Local Feature Matching Revisited
by: Nordström, David, et al.
Published: (2026)
by: Nordström, David, et al.
Published: (2026)
ViT-2SPN: Vision Transformer-based Dual-Stream Self-Supervised Pretraining Networks for Retinal OCT Classification
by: Saraei, Mohammadreza, et al.
Published: (2025)
by: Saraei, Mohammadreza, et al.
Published: (2025)
DeDoDe v2: Analyzing and Improving the DeDoDe Keypoint Detector
by: Edstedt, Johan, et al.
Published: (2024)
by: Edstedt, Johan, et al.
Published: (2024)
Exploring the Synergies of Hybrid CNNs and ViTs Architectures for Computer Vision: A survey
by: Yunusa, Haruna, et al.
Published: (2024)
by: Yunusa, Haruna, et al.
Published: (2024)
Intriguing Frequency Interpretation of Adversarial Robustness for CNNs and ViTs
by: Chen, Lu, et al.
Published: (2025)
by: Chen, Lu, et al.
Published: (2025)
RefTr: Recurrent Refinement of Confluent Trajectories for 3D Vascular Tree Centerlines
by: Naeem, Roman, et al.
Published: (2025)
by: Naeem, Roman, et al.
Published: (2025)
Equi-ViT: Rotational Equivariant Vision Transformer for Robust Histopathology Analysis
by: Chen, Fuyao, et al.
Published: (2026)
by: Chen, Fuyao, et al.
Published: (2026)
Platonic Transformers: A Solid Choice For Equivariance
by: Islam, Mohammad Mohaiminul, et al.
Published: (2025)
by: Islam, Mohammad Mohaiminul, et al.
Published: (2025)
Token Cropr: Faster ViTs for Quite a Few Tasks
by: Bergner, Benjamin, et al.
Published: (2024)
by: Bergner, Benjamin, et al.
Published: (2024)
VIVID-Med: LLM-Supervised Structured Pretraining for Deployable Medical ViTs
by: Wang, Xiyao, et al.
Published: (2026)
by: Wang, Xiyao, et al.
Published: (2026)
CLAMP-ViT: Contrastive Data-Free Learning for Adaptive Post-Training Quantization of ViTs
by: Ramachandran, Akshat, et al.
Published: (2024)
by: Ramachandran, Akshat, et al.
Published: (2024)
RoMa v2: Harder Better Faster Denser Feature Matching
by: Edstedt, Johan, et al.
Published: (2025)
by: Edstedt, Johan, et al.
Published: (2025)
MobilePlantViT: A Mobile-friendly Hybrid ViT for Generalized Plant Disease Image Classification
by: Tonmoy, Moshiur Rahman, et al.
Published: (2025)
by: Tonmoy, Moshiur Rahman, et al.
Published: (2025)
Adapting Pretrained ViTs with Convolution Injector for Visuo-Motor Control
by: Hwang, Dongyoon, et al.
Published: (2024)
by: Hwang, Dongyoon, et al.
Published: (2024)
3D-Consistent Multi-View Editing by Correspondence Guidance
by: Bengtson, Josef, et al.
Published: (2025)
by: Bengtson, Josef, et al.
Published: (2025)
ARTA: Adaptive Mixed-Resolution Token Allocation for Efficient Dense Feature Extraction
by: Hagerman, David, et al.
Published: (2026)
by: Hagerman, David, et al.
Published: (2026)
FasterViT: Fast Vision Transformers with Hierarchical Attention
by: Hatamizadeh, Ali, et al.
Published: (2023)
by: Hatamizadeh, Ali, et al.
Published: (2023)
ViT-ProtoNet for Few-Shot Image Classification: A Multi-Benchmark Evaluation
by: Mutlu, Abdulvahap, et al.
Published: (2025)
by: Mutlu, Abdulvahap, et al.
Published: (2025)
Causality $\neq$ Decodability, and Vice Versa: Lessons from Interpreting Counting ViTs
by: Huang, Lianghuan, et al.
Published: (2025)
by: Huang, Lianghuan, et al.
Published: (2025)
VariViT: A Vision Transformer for Variable Image Sizes
by: Varma, Aswathi, et al.
Published: (2026)
by: Varma, Aswathi, et al.
Published: (2026)
ScriptViT: Vision Transformer-Based Personalized Handwriting Generation
by: Acharya, Sajjan, et al.
Published: (2025)
by: Acharya, Sajjan, et al.
Published: (2025)
DaD: Distilled Reinforcement Learning for Diverse Keypoint Detection
by: Edstedt, Johan, et al.
Published: (2025)
by: Edstedt, Johan, et al.
Published: (2025)
OmniPatch: A Universal Adversarial Patch for ViT-CNN Cross-Architecture Transfer in Semantic Segmentation
by: Aggarwal, Aarush, et al.
Published: (2026)
by: Aggarwal, Aarush, et al.
Published: (2026)
Register and [CLS] tokens yield a decoupling of local and global features in large ViTs
by: Lappe, Alexander, et al.
Published: (2025)
by: Lappe, Alexander, et al.
Published: (2025)
Similar Items
-
MuM: Multi-View Masked Image Modeling for 3D Vision
by: Nordström, David, et al.
Published: (2025) -
Who Handles Orientation? Investigating Invariance in Feature Matching
by: Nordström, David, et al.
Published: (2026) -
Flopping for FLOPs: Leveraging equivariance for computational efficiency
by: Bökman, Georg, et al.
Published: (2025) -
TAP-ViTs: Task-Adaptive Pruning for On-Device Deployment of Vision Transformers
by: Wang, Zhibo, et al.
Published: (2026) -
Steerers: A framework for rotation equivariant keypoint descriptors
by: Bökman, Georg, et al.
Published: (2023)