Diversity-Guided MLP Reduction for Efficient Large Vision Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Shen, Chengchao, Zhu, Hourun, Fang, Gongfan, Wang, Jianxin, Wang, Xinchao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PixelThink: Towards Efficient Chain-of-Pixel Reasoning
by: Wang, Song, et al.
Published: (2025)
by: Wang, Song, et al.
Published: (2025)
Learning Compact Vision Tokens for Efficient Large Multimodal Models
by: Tang, Hao, et al.
Published: (2025)
by: Tang, Hao, et al.
Published: (2025)
Adaptive MLP Pruning for Large Vision Transformers
by: Shen, Chengchao
Published: (2026)
by: Shen, Chengchao
Published: (2026)
SDMPrune: Self-Distillation MLP Pruning for Efficient Large Language Models
by: Zhu, Hourun, et al.
Published: (2025)
by: Zhu, Hourun, et al.
Published: (2025)
Diversity Matters: Revisiting Test-Time Compute in Vision-Language Models
by: Tong, Yijie, et al.
Published: (2026)
by: Tong, Yijie, et al.
Published: (2026)
Parameter-Efficient Subspace Decoupling ViT for Mitigating Multi-Task Negative Transfer in Histological Scoring
by: Huang, Youhan, et al.
Published: (2026)
by: Huang, Youhan, et al.
Published: (2026)
Learning-to-Cache: Accelerating Diffusion Transformer via Layer Caching
by: Ma, Xinyin, et al.
Published: (2024)
by: Ma, Xinyin, et al.
Published: (2024)
Unveiling the Visual Counting Bottleneck in Vision-Language Models
by: Pang, Xingzhou, et al.
Published: (2026)
by: Pang, Xingzhou, et al.
Published: (2026)
TinyFusion: Diffusion Transformers Learned Shallow
by: Fang, Gongfan, et al.
Published: (2024)
by: Fang, Gongfan, et al.
Published: (2024)
TPC-ViT: Token Propagation Controller for Efficient Vision Transformer
by: Zhu, Wentao
Published: (2024)
by: Zhu, Wentao
Published: (2024)
Isomorphic Pruning for Vision Models
by: Fang, Gongfan, et al.
Published: (2024)
by: Fang, Gongfan, et al.
Published: (2024)
Cross-Modal Coordination Across a Diverse Set of Input Modalities
by: Sánchez, Jorge, et al.
Published: (2024)
by: Sánchez, Jorge, et al.
Published: (2024)
Balanced Multi-modal Federated Learning via Cross-Modal Infiltration
by: Fan, Yunfeng, et al.
Published: (2023)
by: Fan, Yunfeng, et al.
Published: (2023)
Zero-shot image privacy classification with Vision-Language Models
by: Baia, Alina Elena, et al.
Published: (2025)
by: Baia, Alina Elena, et al.
Published: (2025)
Text-Guided Image Invariant Feature Learning for Robust Image Watermarking
by: Ahtesham, Muhammad, et al.
Published: (2025)
by: Ahtesham, Muhammad, et al.
Published: (2025)
Efficient Token Compression for Vision Transformer with Spatial Information Preserved
by: Mao, Junzhu, et al.
Published: (2025)
by: Mao, Junzhu, et al.
Published: (2025)
Rethinking Few-Shot Adaptation of Vision-Language Models in Two Stages
by: Farina, Matteo, et al.
Published: (2025)
by: Farina, Matteo, et al.
Published: (2025)
Relating CNN-Transformer Fusion Network for Change Detection
by: Gao, Yuhao, et al.
Published: (2024)
by: Gao, Yuhao, et al.
Published: (2024)
Multimodal Transformer With a Low-Computational-Cost Guarantee
by: Park, Sungjin, et al.
Published: (2024)
by: Park, Sungjin, et al.
Published: (2024)
Detecting Content Rating Violations in Android Applications: A Vision-Language Approach
by: Denipitiyage, D., et al.
Published: (2025)
by: Denipitiyage, D., et al.
Published: (2025)
CreativeVR: Diffusion-Prior-Guided Approach for Structure and Motion Restoration in Generative and Real Videos
by: Panambur, Tejas, et al.
Published: (2025)
by: Panambur, Tejas, et al.
Published: (2025)
EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoE
by: Chen, Junyi, et al.
Published: (2023)
by: Chen, Junyi, et al.
Published: (2023)
Improving Accuracy and Generalization for Efficient Visual Tracking
by: Zaveri, Ram, et al.
Published: (2024)
by: Zaveri, Ram, et al.
Published: (2024)
T-TAME: Trainable Attention Mechanism for Explaining Convolutional Networks and Vision Transformers
by: Ntrougkas, Mariano V., et al.
Published: (2024)
by: Ntrougkas, Mariano V., et al.
Published: (2024)
Multi-Modal Image Fusion via Intervention-Stable Feature Learning
by: Wang, Xue, et al.
Published: (2026)
by: Wang, Xue, et al.
Published: (2026)
Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided Decoding
by: Deng, Ailin, et al.
Published: (2024)
by: Deng, Ailin, et al.
Published: (2024)
LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos
by: Geng, Tiantian, et al.
Published: (2024)
by: Geng, Tiantian, et al.
Published: (2024)
Multimodal Learning on Low-Quality Data with Conformal Predictive Self-Calibration
by: Jiang, Xun, et al.
Published: (2026)
by: Jiang, Xun, et al.
Published: (2026)
Bridging Compressed Image Latents and Multimodal Large Language Models
by: Kao, Chia-Hao, et al.
Published: (2024)
by: Kao, Chia-Hao, et al.
Published: (2024)
4D Multimodal Co-attention Fusion Network with Latent Contrastive Alignment for Alzheimer's Diagnosis
by: Wei, Yuxiang, et al.
Published: (2025)
by: Wei, Yuxiang, et al.
Published: (2025)
LinVT: Empower Your Image-level Large Language Model to Understand Videos
by: Gao, Lishuai, et al.
Published: (2024)
by: Gao, Lishuai, et al.
Published: (2024)
Knowledge Bridger: Towards Training-free Missing Modality Completion
by: Ke, Guanzhou, et al.
Published: (2025)
by: Ke, Guanzhou, et al.
Published: (2025)
PMPGuard: Catching Pseudo-Matched Pairs in Remote Sensing Image-Text Retrieval
by: Ouyang, Pengxiang, et al.
Published: (2025)
by: Ouyang, Pengxiang, et al.
Published: (2025)
Decoupling Spatio-Temporal Adapter for Fine-Grained Badminton Action Localization
by: Wang, Tianyu, et al.
Published: (2026)
by: Wang, Tianyu, et al.
Published: (2026)
Semantic-Aware Adversarial Training for Reliable Deep Hashing Retrieval
by: Yuan, Xu, et al.
Published: (2023)
by: Yuan, Xu, et al.
Published: (2023)
Embracing Collaboration Over Competition: Condensing Multiple Prompts for Visual In-Context Learning
by: Wang, Jinpeng, et al.
Published: (2025)
by: Wang, Jinpeng, et al.
Published: (2025)
Residual Prior-driven Frequency-aware Network for Image Fusion
by: Zheng, Guan, et al.
Published: (2025)
by: Zheng, Guan, et al.
Published: (2025)
Regularized Contrastive Partial Multi-view Outlier Detection
by: Wang, Yijia, et al.
Published: (2024)
by: Wang, Yijia, et al.
Published: (2024)
360VFI: A Dataset and Benchmark for Omnidirectional Video Frame Interpolation
by: Lu, Wenxuan, et al.
Published: (2024)
by: Lu, Wenxuan, et al.
Published: (2024)
RMAdapter: Reconstruction-based Multi-Modal Adapter for Vision-Language Models
by: Lin, Xiang, et al.
Published: (2025)
by: Lin, Xiang, et al.
Published: (2025)
Similar Items
-
PixelThink: Towards Efficient Chain-of-Pixel Reasoning
by: Wang, Song, et al.
Published: (2025) -
Learning Compact Vision Tokens for Efficient Large Multimodal Models
by: Tang, Hao, et al.
Published: (2025) -
Adaptive MLP Pruning for Large Vision Transformers
by: Shen, Chengchao
Published: (2026) -
SDMPrune: Self-Distillation MLP Pruning for Efficient Large Language Models
by: Zhu, Hourun, et al.
Published: (2025) -
Diversity Matters: Revisiting Test-Time Compute in Vision-Language Models
by: Tong, Yijie, et al.
Published: (2026)