You Only Need Less Attention at Each Stage in Vision Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Shuoxi, Liu, Hanpeng, Lin, Stephen, He, Kun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Separators in Enhancing Autoregressive Pretraining for Vision Mamba
von: Liu, Hanpeng, et al.
Veröffentlicht: (2026)
von: Liu, Hanpeng, et al.
Veröffentlicht: (2026)
Siamese Transformer Networks for Few-shot Image Classification
von: Jiang, Weihao, et al.
Veröffentlicht: (2024)
von: Jiang, Weihao, et al.
Veröffentlicht: (2024)
iGVLM: Dynamic Instruction-Guided Vision Encoding for Question-Aware Multimodal Understanding
von: Liu, Hanpeng, et al.
Veröffentlicht: (2026)
von: Liu, Hanpeng, et al.
Veröffentlicht: (2026)
Neural Collapse Inspired Knowledge Distillation
von: Zhang, Shuoxi, et al.
Veröffentlicht: (2024)
von: Zhang, Shuoxi, et al.
Veröffentlicht: (2024)
ITO: Images and Texts as One via Synergizing Multiple Alignment and Training-Time Fusion
von: Liu, Hanpeng, et al.
Veröffentlicht: (2026)
von: Liu, Hanpeng, et al.
Veröffentlicht: (2026)
Intra-task Mutual Attention based Vision Transformer for Few-Shot Learning
von: Jiang, Weihao, et al.
Veröffentlicht: (2024)
von: Jiang, Weihao, et al.
Veröffentlicht: (2024)
You Only Need One Stage: Novel-View Synthesis From A Single Blind Face Image
von: Wang, Taoyue, et al.
Veröffentlicht: (2026)
von: Wang, Taoyue, et al.
Veröffentlicht: (2026)
AgenticOCR: Parsing Only What You Need for Efficient Retrieval-Augmented Generation
von: Wang, Zhengren, et al.
Veröffentlicht: (2026)
von: Wang, Zhengren, et al.
Veröffentlicht: (2026)
You Only Need Two Detectors to Achieve Multi-Modal 3D Multi-Object Tracking
von: Wang, Xiyang, et al.
Veröffentlicht: (2023)
von: Wang, Xiyang, et al.
Veröffentlicht: (2023)
Spiking Transformer:Introducing Accurate Addition-Only Spiking Self-Attention for Transformer
von: Guo, Yufei, et al.
Veröffentlicht: (2025)
von: Guo, Yufei, et al.
Veröffentlicht: (2025)
CAMixerSR: Only Details Need More "Attention"
von: Wang, Yan, et al.
Veröffentlicht: (2024)
von: Wang, Yan, et al.
Veröffentlicht: (2024)
You Only Need Half: Boosting Data Augmentation by Using Partial Content
von: Hu, Juntao, et al.
Veröffentlicht: (2024)
von: Hu, Juntao, et al.
Veröffentlicht: (2024)
Attention Is All You Need For Mixture-of-Depths Routing
von: Gadhikar, Advait, et al.
Veröffentlicht: (2024)
von: Gadhikar, Advait, et al.
Veröffentlicht: (2024)
Vision Transformers with Hierarchical Attention
von: Liu, Yun, et al.
Veröffentlicht: (2021)
von: Liu, Yun, et al.
Veröffentlicht: (2021)
Attn-Adapter: Attention Is All You Need for Online Few-shot Learner of Vision-Language Model
von: Bui, Phuoc-Nguyen, et al.
Veröffentlicht: (2025)
von: Bui, Phuoc-Nguyen, et al.
Veröffentlicht: (2025)
ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention
von: Liu, Wenjie, et al.
Veröffentlicht: (2026)
von: Liu, Wenjie, et al.
Veröffentlicht: (2026)
Take Only What You Need: Rank Minimization as an Implicit Forgetting Regularizer in Continual Learning
von: Lu, Haodong, et al.
Veröffentlicht: (2024)
von: Lu, Haodong, et al.
Veröffentlicht: (2024)
YoNoSplat: You Only Need One Model for Feedforward 3D Gaussian Splatting
von: Ye, Botao, et al.
Veröffentlicht: (2025)
von: Ye, Botao, et al.
Veröffentlicht: (2025)
Text-Guided Attention is All You Need for Zero-Shot Robustness in Vision-Language Models
von: Yu, Lu, et al.
Veröffentlicht: (2024)
von: Yu, Lu, et al.
Veröffentlicht: (2024)
Polyline Path Masked Attention for Vision Transformer
von: Zhao, Zhongchen, et al.
Veröffentlicht: (2025)
von: Zhao, Zhongchen, et al.
Veröffentlicht: (2025)
Ideal Registration? Segmentation is All You Need
von: Chen, Xiang, et al.
Veröffentlicht: (2025)
von: Chen, Xiang, et al.
Veröffentlicht: (2025)
SeTformer is What You Need for Vision and Language
von: Shamsolmoali, Pourya, et al.
Veröffentlicht: (2024)
von: Shamsolmoali, Pourya, et al.
Veröffentlicht: (2024)
Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding
von: Kang, Seil, et al.
Veröffentlicht: (2025)
von: Kang, Seil, et al.
Veröffentlicht: (2025)
Representative Attention For Vision Transformers
von: Li, Yuntong, et al.
Veröffentlicht: (2026)
von: Li, Yuntong, et al.
Veröffentlicht: (2026)
You Only Need One Step: Fast Super-Resolution with Stable Diffusion via Scale Distillation
von: Noroozi, Mehdi, et al.
Veröffentlicht: (2024)
von: Noroozi, Mehdi, et al.
Veröffentlicht: (2024)
MaTe: Images Are All You Need for Material Transfer via Diffusion Transformer
von: Huang, Nisha, et al.
Veröffentlicht: (2026)
von: Huang, Nisha, et al.
Veröffentlicht: (2026)
Vision Transformers Need Registers
von: Darcet, Timothée, et al.
Veröffentlicht: (2023)
von: Darcet, Timothée, et al.
Veröffentlicht: (2023)
Unlearnable 3D Point Clouds: Class-wise Transformation Is All You Need
von: Wang, Xianlong, et al.
Veröffentlicht: (2024)
von: Wang, Xianlong, et al.
Veröffentlicht: (2024)
Autoregressive Image Generation Needs Only a Few Lines of Cached Tokens
von: Qin, Ziran, et al.
Veröffentlicht: (2025)
von: Qin, Ziran, et al.
Veröffentlicht: (2025)
You Only Need One Color Space: An Efficient Network for Low-light Image Enhancement
von: Yan, Qingsen, et al.
Veröffentlicht: (2024)
von: Yan, Qingsen, et al.
Veröffentlicht: (2024)
You Only Speak Once to See
von: Yang, Wenhao, et al.
Veröffentlicht: (2024)
von: Yang, Wenhao, et al.
Veröffentlicht: (2024)
Catch-Up Distillation: You Only Need to Train Once for Accelerating Sampling
von: Shao, Shitong, et al.
Veröffentlicht: (2023)
von: Shao, Shitong, et al.
Veröffentlicht: (2023)
Transferable-guided Attention Is All You Need for Video Domain Adaptation
von: Sacilotti, André, et al.
Veröffentlicht: (2024)
von: Sacilotti, André, et al.
Veröffentlicht: (2024)
BinaryAttention: One-Bit QK-Attention for Vision and Diffusion Transformers
von: Xiao, Chaodong, et al.
Veröffentlicht: (2026)
von: Xiao, Chaodong, et al.
Veröffentlicht: (2026)
Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
YotoR-You Only Transform One Representation
von: Villa, José Ignacio Díaz, et al.
Veröffentlicht: (2024)
von: Villa, José Ignacio Díaz, et al.
Veröffentlicht: (2024)
Calibration Attention: Learning Reliability-Aware Representations for Vision Transformers
von: Liang, Wenhao, et al.
Veröffentlicht: (2025)
von: Liang, Wenhao, et al.
Veröffentlicht: (2025)
You Only Learn One Query: Learning Unified Human Query for Single-Stage Multi-Person Multi-Task Human-Centric Perception
von: Jin, Sheng, et al.
Veröffentlicht: (2023)
von: Jin, Sheng, et al.
Veröffentlicht: (2023)
Vision Also You Need: Navigating Out-of-Distribution Detection with Multimodal Large Language Model
von: Xu, Haoran, et al.
Veröffentlicht: (2026)
von: Xu, Haoran, et al.
Veröffentlicht: (2026)
ECViT: Efficient Convolutional Vision Transformer with Local-Attention and Multi-scale Stages
von: Qian, Zhoujie
Veröffentlicht: (2025)
von: Qian, Zhoujie
Veröffentlicht: (2025)
Ähnliche Einträge
-
Separators in Enhancing Autoregressive Pretraining for Vision Mamba
von: Liu, Hanpeng, et al.
Veröffentlicht: (2026) -
Siamese Transformer Networks for Few-shot Image Classification
von: Jiang, Weihao, et al.
Veröffentlicht: (2024) -
iGVLM: Dynamic Instruction-Guided Vision Encoding for Question-Aware Multimodal Understanding
von: Liu, Hanpeng, et al.
Veröffentlicht: (2026) -
Neural Collapse Inspired Knowledge Distillation
von: Zhang, Shuoxi, et al.
Veröffentlicht: (2024) -
ITO: Images and Texts as One via Synergizing Multiple Alignment and Training-Time Fusion
von: Liu, Hanpeng, et al.
Veröffentlicht: (2026)