Multi-Tailed Vision Transformer for Efficient Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yunke, Du, Bo, Wang, Wenyuan, Xu, Chang |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MambaScope: Coarse-to-Fine Scoping for Efficient Vision Mamba
by: Liu, Shanhui, et al.
Published: (2025)
by: Liu, Shanhui, et al.
Published: (2025)
Rethinking Visual Token Reduction in LVLMs Under Cross-Modal Misalignment
by: Xu, Rui, et al.
Published: (2025)
by: Xu, Rui, et al.
Published: (2025)
UniV2D: Bridging Visual Restoration and Semantic Perception for Underwater Salient Object Detection
by: Chang, Laibin, et al.
Published: (2026)
by: Chang, Laibin, et al.
Published: (2026)
MAEDiff: Masked Autoencoder-enhanced Diffusion Models for Unsupervised Anomaly Detection in Brain Images
by: Xu, Rui, et al.
Published: (2024)
by: Xu, Rui, et al.
Published: (2024)
CollagePrompt: A Benchmark for Budget-Friendly Visual Recognition with GPT-4V
by: Xu, Siyu, et al.
Published: (2024)
by: Xu, Siyu, et al.
Published: (2024)
Marine Saliency Segmenter: Object-Focused Conditional Diffusion with Region-Level Semantic Knowledge Distillation
by: Chang, Laibin, et al.
Published: (2025)
by: Chang, Laibin, et al.
Published: (2025)
Visual Imitation Learning with Calibrated Contrastive Representation
by: Wang, Yunke, et al.
Published: (2024)
by: Wang, Yunke, et al.
Published: (2024)
Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation
by: Hui, Chenyu, et al.
Published: (2026)
by: Hui, Chenyu, et al.
Published: (2026)
BEAT: Rhythm-Elastic Alignment for Agentic Music-guided Movie Trailer Generation
by: Wang, Yutong, et al.
Published: (2026)
by: Wang, Yutong, et al.
Published: (2026)
VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching
by: Xu, Siyu, et al.
Published: (2025)
by: Xu, Siyu, et al.
Published: (2025)
See What Matters: Differentiable Grid Sample Pruning for Generalizable Vision-Language-Action Model
by: Feng, Yixu, et al.
Published: (2026)
by: Feng, Yixu, et al.
Published: (2026)
PARE: Pruning and Adaptive Routing for Efficient Video Generation
by: Wang, Yutong, et al.
Published: (2026)
by: Wang, Yutong, et al.
Published: (2026)
Depth-Wise Convolutions in Vision Transformers for Efficient Training on Small Datasets
by: Zhang, Tianxiao, et al.
Published: (2024)
by: Zhang, Tianxiao, et al.
Published: (2024)
Sparse-Tuning: Adapting Vision Transformers with Efficient Fine-tuning and Inference
by: Liu, Ting, et al.
Published: (2024)
by: Liu, Ting, et al.
Published: (2024)
LeMeViT: Efficient Vision Transformer with Learnable Meta Tokens for Remote Sensing Image Interpretation
by: Jiang, Wentao, et al.
Published: (2024)
by: Jiang, Wentao, et al.
Published: (2024)
Early Semantic Grounding in Image Editing Models for Zero-Shot Referring Image Segmentation
by: He, Jingxuan, et al.
Published: (2026)
by: He, Jingxuan, et al.
Published: (2026)
Transforming Vision Transformer: Towards Efficient Multi-Task Asynchronous Learning
by: Zhong, Hanwen, et al.
Published: (2025)
by: Zhong, Hanwen, et al.
Published: (2025)
Improving Vision Transformers by Overlapping Heads in Multi-Head Self-Attention
by: Zhang, Tianxiao, et al.
Published: (2024)
by: Zhang, Tianxiao, et al.
Published: (2024)
Empowering Vision Transformers with Multi-Scale Causal Intervention for Long-Tailed Image Classification
by: Yan, Xiaoshuo, et al.
Published: (2025)
by: Yan, Xiaoshuo, et al.
Published: (2025)
Rethinking Where to Edit: Task-Aware Localization for Instruction-Based Image Editing
by: He, Jingxuan, et al.
Published: (2026)
by: He, Jingxuan, et al.
Published: (2026)
FusionSAM: Visual Multi-Modal Learning with Segment Anything
by: Li, Daixun, et al.
Published: (2024)
by: Li, Daixun, et al.
Published: (2024)
Multiple-Exit Tuning: Towards Inference-Efficient Adaptation for Vision Transformer
by: Liu, Zheng, et al.
Published: (2024)
by: Liu, Zheng, et al.
Published: (2024)
Slicing Vision Transformer for Flexible Inference
by: Zhang, Yitian, et al.
Published: (2024)
by: Zhang, Yitian, et al.
Published: (2024)
Cross-Self KV Cache Pruning for Efficient Vision-Language Inference
by: Pei, Xiaohuan, et al.
Published: (2024)
by: Pei, Xiaohuan, et al.
Published: (2024)
Multi-Attribute Vision Transformers are Efficient and Robust Learners
by: Gani, Hanan, et al.
Published: (2024)
by: Gani, Hanan, et al.
Published: (2024)
GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision Transformer
by: Jia, Ding, et al.
Published: (2024)
by: Jia, Ding, et al.
Published: (2024)
IML-ViT: Benchmarking Image Manipulation Localization by Vision Transformer
by: Ma, Xiaochen, et al.
Published: (2023)
by: Ma, Xiaochen, et al.
Published: (2023)
ViTCN: Vision Transformer Contrastive Network For Reasoning
by: Song, Bo, et al.
Published: (2024)
by: Song, Bo, et al.
Published: (2024)
PPT: Token Pruning and Pooling for Efficient Vision Transformers
by: Wu, Xinjian, et al.
Published: (2023)
by: Wu, Xinjian, et al.
Published: (2023)
Parameter-Efficient and Memory-Efficient Tuning for Vision Transformer: A Disentangled Approach
by: Zhang, Taolin, et al.
Published: (2024)
by: Zhang, Taolin, et al.
Published: (2024)
big.LITTLE Vision Transformer for Efficient Visual Recognition
by: Guo, He, et al.
Published: (2024)
by: Guo, He, et al.
Published: (2024)
Efficient Inference of Vision Instruction-Following Models with Elastic Cache
by: Liu, Zuyan, et al.
Published: (2024)
by: Liu, Zuyan, et al.
Published: (2024)
Protego: Detecting Adversarial Examples for Vision Transformers via Intrinsic Capabilities
by: Wu, Jialin, et al.
Published: (2025)
by: Wu, Jialin, et al.
Published: (2025)
EA-ViT: Efficient Adaptation for Elastic Vision Transformer
by: Zhu, Chen, et al.
Published: (2025)
by: Zhu, Chen, et al.
Published: (2025)
Efficient Partitioning Vision Transformer on Edge Devices for Distributed Inference
by: Liu, Xiang, et al.
Published: (2024)
by: Liu, Xiang, et al.
Published: (2024)
HSViT: Horizontally Scalable Vision Transformer
by: Xu, Chenhao, et al.
Published: (2024)
by: Xu, Chenhao, et al.
Published: (2024)
Token Compensator: Altering Inference Cost of Vision Transformer without Re-Tuning
by: Jie, Shibo, et al.
Published: (2024)
by: Jie, Shibo, et al.
Published: (2024)
Block-based Symmetric Pruning and Fusion for Efficient Vision Transformers
by: Hsieh, Yi-Kuan, et al.
Published: (2025)
by: Hsieh, Yi-Kuan, et al.
Published: (2025)
Modest-Align: Data-Efficient Alignment for Vision-Language Models
by: Liu, Jiaxiang, et al.
Published: (2025)
by: Liu, Jiaxiang, et al.
Published: (2025)
Artifacts and Attention Sinks: Structured Approximations for Efficient Vision Transformers
by: Lu, Andrew, et al.
Published: (2025)
by: Lu, Andrew, et al.
Published: (2025)
Similar Items
-
MambaScope: Coarse-to-Fine Scoping for Efficient Vision Mamba
by: Liu, Shanhui, et al.
Published: (2025) -
Rethinking Visual Token Reduction in LVLMs Under Cross-Modal Misalignment
by: Xu, Rui, et al.
Published: (2025) -
UniV2D: Bridging Visual Restoration and Semantic Perception for Underwater Salient Object Detection
by: Chang, Laibin, et al.
Published: (2026) -
MAEDiff: Masked Autoencoder-enhanced Diffusion Models for Unsupervised Anomaly Detection in Brain Images
by: Xu, Rui, et al.
Published: (2024) -
CollagePrompt: A Benchmark for Budget-Friendly Visual Recognition with GPT-4V
by: Xu, Siyu, et al.
Published: (2024)