ViT$^3$: Unlocking Test-Time Training in Vision
Fuente:
arXiv
Salvato in:
| Autori principali: | Han, Dongchen, Li, Yining, Li, Tianyu, Cao, Zixuan, Wang, Ziming, Song, Jun, Cheng, Yu, Zheng, Bo, Huang, Gao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VAT: Vision Action Transformer by Unlocking Full Representation of ViT
di: Li, Wenhao, et al.
Pubblicazione: (2025)
di: Li, Wenhao, et al.
Pubblicazione: (2025)
Vision Transformers are Circulant Attention Learners
di: Han, Dongchen, et al.
Pubblicazione: (2025)
di: Han, Dongchen, et al.
Pubblicazione: (2025)
MPTQ-ViT: Mixed-Precision Post-Training Quantization for Vision Transformer
di: Tai, Yu-Shan, et al.
Pubblicazione: (2024)
di: Tai, Yu-Shan, et al.
Pubblicazione: (2024)
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs
di: Kuzucu, Selim, et al.
Pubblicazione: (2025)
di: Kuzucu, Selim, et al.
Pubblicazione: (2025)
ADFQ-ViT: Activation-Distribution-Friendly Post-Training Quantization for Vision Transformers
di: Jiang, Yanfeng, et al.
Pubblicazione: (2024)
di: Jiang, Yanfeng, et al.
Pubblicazione: (2024)
LF-ViT: Reducing Spatial Redundancy in Vision Transformer for Efficient Image Recognition
di: Hu, Youbing, et al.
Pubblicazione: (2024)
di: Hu, Youbing, et al.
Pubblicazione: (2024)
ViT-AdaLA: Adapting Vision Transformers with Linear Attention
di: Li, Yifan, et al.
Pubblicazione: (2026)
di: Li, Yifan, et al.
Pubblicazione: (2026)
Modulating CNN Features with Pre-Trained ViT Representations for Open-Vocabulary Object Detection
di: Gao, Xiangyu, et al.
Pubblicazione: (2025)
di: Gao, Xiangyu, et al.
Pubblicazione: (2025)
SAC-ViT: Semantic-Aware Clustering Vision Transformer with Early Exit
di: Hu, Youbing, et al.
Pubblicazione: (2025)
di: Hu, Youbing, et al.
Pubblicazione: (2025)
I&S-ViT: An Inclusive & Stable Method for Pushing the Limit of Post-Training ViTs Quantization
di: Zhong, Yunshan, et al.
Pubblicazione: (2023)
di: Zhong, Yunshan, et al.
Pubblicazione: (2023)
Dynamic Tuning Towards Parameter and Inference Efficiency for ViT Adaptation
di: Zhao, Wangbo, et al.
Pubblicazione: (2024)
di: Zhao, Wangbo, et al.
Pubblicazione: (2024)
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads
di: Li, Yifan, et al.
Pubblicazione: (2025)
di: Li, Yifan, et al.
Pubblicazione: (2025)
IML-ViT: Benchmarking Image Manipulation Localization by Vision Transformer
di: Ma, Xiaochen, et al.
Pubblicazione: (2023)
di: Ma, Xiaochen, et al.
Pubblicazione: (2023)
HIRI-ViT: Scaling Vision Transformer with High Resolution Inputs
di: Yao, Ting, et al.
Pubblicazione: (2024)
di: Yao, Ting, et al.
Pubblicazione: (2024)
ViT-5: Vision Transformers for The Mid-2020s
di: Wang, Feng, et al.
Pubblicazione: (2026)
di: Wang, Feng, et al.
Pubblicazione: (2026)
APHQ-ViT: Post-Training Quantization with Average Perturbation Hessian Based Reconstruction for Vision Transformers
di: Wu, Zhuguanyu, et al.
Pubblicazione: (2025)
di: Wu, Zhuguanyu, et al.
Pubblicazione: (2025)
Intriguing Frequency Interpretation of Adversarial Robustness for CNNs and ViTs
di: Chen, Lu, et al.
Pubblicazione: (2025)
di: Chen, Lu, et al.
Pubblicazione: (2025)
ViT-Linearizer: Distilling Quadratic Knowledge into Linear-Time Vision Models
di: Wei, Guoyizhe, et al.
Pubblicazione: (2025)
di: Wei, Guoyizhe, et al.
Pubblicazione: (2025)
CLAMP-ViT: Contrastive Data-Free Learning for Adaptive Post-Training Quantization of ViTs
di: Ramachandran, Akshat, et al.
Pubblicazione: (2024)
di: Ramachandran, Akshat, et al.
Pubblicazione: (2024)
Demystify Mamba in Vision: A Linear Attention Perspective
di: Han, Dongchen, et al.
Pubblicazione: (2024)
di: Han, Dongchen, et al.
Pubblicazione: (2024)
ACC-ViT : Atrous Convolution's Comeback in Vision Transformers
di: Ibtehaz, Nabil, et al.
Pubblicazione: (2024)
di: Ibtehaz, Nabil, et al.
Pubblicazione: (2024)
EA-ViT: Efficient Adaptation for Elastic Vision Transformer
di: Zhu, Chen, et al.
Pubblicazione: (2025)
di: Zhu, Chen, et al.
Pubblicazione: (2025)
RepViT: Revisiting Mobile CNN From ViT Perspective
di: Wang, Ao, et al.
Pubblicazione: (2023)
di: Wang, Ao, et al.
Pubblicazione: (2023)
Rethinking Random Masking in Self-Distillation on ViT
di: Seong, Jihyeon, et al.
Pubblicazione: (2025)
di: Seong, Jihyeon, et al.
Pubblicazione: (2025)
Linear-Time Global Visual Modeling without Explicit Attention
di: He, Ruize, et al.
Pubblicazione: (2026)
di: He, Ruize, et al.
Pubblicazione: (2026)
Trio-ViT: Post-Training Quantization and Acceleration for Softmax-Free Efficient Vision Transformer
di: Shi, Huihong, et al.
Pubblicazione: (2024)
di: Shi, Huihong, et al.
Pubblicazione: (2024)
SegDebias: Test-Time Bias Mitigation for ViT-Based CLIP via Segmentation
di: Wu, Fangyu, et al.
Pubblicazione: (2025)
di: Wu, Fangyu, et al.
Pubblicazione: (2025)
FastPose-ViT: A Vision Transformer for Real-Time Spacecraft Pose Estimation
di: Ancey, Pierre, et al.
Pubblicazione: (2025)
di: Ancey, Pierre, et al.
Pubblicazione: (2025)
ViT-Lens: Towards Omni-modal Representations
di: Lei, Weixian, et al.
Pubblicazione: (2023)
di: Lei, Weixian, et al.
Pubblicazione: (2023)
ViT-Explainer: An Interactive Walkthrough of the Vision Transformer Pipeline
di: Hernandez, Juan Manuel, et al.
Pubblicazione: (2026)
di: Hernandez, Juan Manuel, et al.
Pubblicazione: (2026)
ViT-EnsembleAttack: Augmenting Ensemble Models for Stronger Adversarial Transferability in Vision Transformers
di: Cao, Hanwen, et al.
Pubblicazione: (2025)
di: Cao, Hanwen, et al.
Pubblicazione: (2025)
GTP-ViT: Efficient Vision Transformers via Graph-based Token Propagation
di: Xu, Xuwei, et al.
Pubblicazione: (2023)
di: Xu, Xuwei, et al.
Pubblicazione: (2023)
IPTQ-ViT: Post-Training Quantization of Non-linear Functions for Integer-only Vision Transformers
di: Kim, Gihwan, et al.
Pubblicazione: (2025)
di: Kim, Gihwan, et al.
Pubblicazione: (2025)
A Hybrid Framework Bridging CNN and ViT based on Theory of Evidence for Diabetic Retinopathy Grading
di: Qiu, Junlai, et al.
Pubblicazione: (2025)
di: Qiu, Junlai, et al.
Pubblicazione: (2025)
Deeper Inside Deep ViT
di: Hong, Sungrae
Pubblicazione: (2025)
di: Hong, Sungrae
Pubblicazione: (2025)
DopQ-ViT: Towards Distribution-Friendly and Outlier-Aware Post-Training Quantization for Vision Transformers
di: Yang, Lianwei, et al.
Pubblicazione: (2024)
di: Yang, Lianwei, et al.
Pubblicazione: (2024)
DeNAS-ViT: Data Efficient NAS-Optimized Vision Transformer for Ultrasound Image Segmentation
di: Chen, Renqi, et al.
Pubblicazione: (2024)
di: Chen, Renqi, et al.
Pubblicazione: (2024)
CAS-ViT: Convolutional Additive Self-attention Vision Transformers for Efficient Mobile Applications
di: Zhang, Tianfang, et al.
Pubblicazione: (2024)
di: Zhang, Tianfang, et al.
Pubblicazione: (2024)
ViT-FIQA: Assessing Face Image Quality using Vision Transformers
di: Atzori, Andrea, et al.
Pubblicazione: (2025)
di: Atzori, Andrea, et al.
Pubblicazione: (2025)
Colinearity Decay: Training Quantization-Friendly ViTs with Outlier Decay
di: Tong, Jin, et al.
Pubblicazione: (2026)
di: Tong, Jin, et al.
Pubblicazione: (2026)
Documenti analoghi
-
VAT: Vision Action Transformer by Unlocking Full Representation of ViT
di: Li, Wenhao, et al.
Pubblicazione: (2025) -
Vision Transformers are Circulant Attention Learners
di: Han, Dongchen, et al.
Pubblicazione: (2025) -
MPTQ-ViT: Mixed-Precision Post-Training Quantization for Vision Transformer
di: Tai, Yu-Shan, et al.
Pubblicazione: (2024) -
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs
di: Kuzucu, Selim, et al.
Pubblicazione: (2025) -
ADFQ-ViT: Activation-Distribution-Friendly Post-Training Quantization for Vision Transformers
di: Jiang, Yanfeng, et al.
Pubblicazione: (2024)