An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Liang, Zhao, Haozhe, Liu, Tianyu, Bai, Shuai, Lin, Junyang, Zhou, Chang, Chang, Baobao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Spark of Vision-Language Intelligence: 2-Dimensional Autoregressive Transformer for Efficient Finegrained Image Generation
von: Chen, Liang, et al.
Veröffentlicht: (2024)
von: Chen, Liang, et al.
Veröffentlicht: (2024)
Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think
von: Chen, Liang, et al.
Veröffentlicht: (2025)
von: Chen, Liang, et al.
Veröffentlicht: (2025)
Looking Beyond Text: Reducing Language bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance
von: Zhao, Haozhe, et al.
Veröffentlicht: (2024)
von: Zhao, Haozhe, et al.
Veröffentlicht: (2024)
VLCache: Computing 2% Vision Tokens and Reusing 98% for Vision-Language Inference
von: Qin, Shengling, et al.
Veröffentlicht: (2025)
von: Qin, Shengling, et al.
Veröffentlicht: (2025)
Turbo: Informativity-Driven Acceleration Plug-In for Vision-Language Large Models
von: Ju, Chen, et al.
Veröffentlicht: (2024)
von: Ju, Chen, et al.
Veröffentlicht: (2024)
Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models
von: Liu, Xuyang, et al.
Veröffentlicht: (2025)
von: Liu, Xuyang, et al.
Veröffentlicht: (2025)
Vote&Mix: Plug-and-Play Token Reduction for Efficient Vision Transformer
von: Peng, Shuai, et al.
Veröffentlicht: (2024)
von: Peng, Shuai, et al.
Veröffentlicht: (2024)
PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain
von: Chen, Liang, et al.
Veröffentlicht: (2024)
von: Chen, Liang, et al.
Veröffentlicht: (2024)
Rethinking Semantic Parsing for Large Language Models: Enhancing LLM Performance with Semantic Hints
von: An, Kaikai, et al.
Veröffentlicht: (2024)
von: An, Kaikai, et al.
Veröffentlicht: (2024)
Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Models
von: Liu, Xuyang, et al.
Veröffentlicht: (2025)
von: Liu, Xuyang, et al.
Veröffentlicht: (2025)
Mitigating Language-Level Performance Disparity in mPLMs via Teacher Language Selection and Cross-lingual Self-Distillation
von: Zhao, Haozhe, et al.
Veröffentlicht: (2024)
von: Zhao, Haozhe, et al.
Veröffentlicht: (2024)
Revisiting Multimodal Positional Encoding in Vision-Language Models
von: Huang, Jie, et al.
Veröffentlicht: (2025)
von: Huang, Jie, et al.
Veröffentlicht: (2025)
Prune2Drive: A Plug-and-Play Framework for Accelerating Vision-Language Models in Autonomous Driving
von: Xiong, Minhao, et al.
Veröffentlicht: (2025)
von: Xiong, Minhao, et al.
Veröffentlicht: (2025)
G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning
von: Chen, Liang, et al.
Veröffentlicht: (2025)
von: Chen, Liang, et al.
Veröffentlicht: (2025)
Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models
von: Chen, Jiaxing, et al.
Veröffentlicht: (2024)
von: Chen, Jiaxing, et al.
Veröffentlicht: (2024)
LLM-KT: Aligning Large Language Models with Knowledge Tracing using a Plug-and-Play Instruction
von: Wang, Ziwei, et al.
Veröffentlicht: (2025)
von: Wang, Ziwei, et al.
Veröffentlicht: (2025)
MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
von: Zhao, Haozhe, et al.
Veröffentlicht: (2023)
von: Zhao, Haozhe, et al.
Veröffentlicht: (2023)
V2Flow: Unifying Visual Tokenization and Large Language Model Vocabularies for Autoregressive Image Generation
von: Zhang, Guiwei, et al.
Veröffentlicht: (2025)
von: Zhang, Guiwei, et al.
Veröffentlicht: (2025)
Can Large Language Models Always Solve Easy Problems if They Can Solve Harder Ones?
von: Yang, Zhe, et al.
Veröffentlicht: (2024)
von: Yang, Zhe, et al.
Veröffentlicht: (2024)
Score-Based Turbo Message Passing for Plug-and-Play Compressive Imaging
von: Cai, Chang, et al.
Veröffentlicht: (2025)
von: Cai, Chang, et al.
Veröffentlicht: (2025)
DynaMo: Accelerating Language Model Inference with Dynamic Multi-Token Sampling
von: Tuli, Shikhar, et al.
Veröffentlicht: (2024)
von: Tuli, Shikhar, et al.
Veröffentlicht: (2024)
An Image is Worth 32 Tokens for Reconstruction and Generation
von: Yu, Qihang, et al.
Veröffentlicht: (2024)
von: Yu, Qihang, et al.
Veröffentlicht: (2024)
Variator: Accelerating Pre-trained Models with Plug-and-Play Compression Modules
von: Xiao, Chaojun, et al.
Veröffentlicht: (2023)
von: Xiao, Chaojun, et al.
Veröffentlicht: (2023)
Memory Decoder: A Pretrained, Plug-and-Play Memory for Large Language Models
von: Cao, Jiaqi, et al.
Veröffentlicht: (2025)
von: Cao, Jiaqi, et al.
Veröffentlicht: (2025)
Aligning Large Language Models to Follow Instructions and Hallucinate Less via Effective Data Filtering
von: Si, Shuzheng, et al.
Veröffentlicht: (2025)
von: Si, Shuzheng, et al.
Veröffentlicht: (2025)
UltraEdit: Instruction-based Fine-Grained Image Editing at Scale
von: Zhao, Haozhe, et al.
Veröffentlicht: (2024)
von: Zhao, Haozhe, et al.
Veröffentlicht: (2024)
Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
von: Wang, Peng, et al.
Veröffentlicht: (2024)
von: Wang, Peng, et al.
Veröffentlicht: (2024)
Score-Based Turbo Message Passing for Plug-and-Play Compressive Image Recovery
von: Cai, Chang, et al.
Veröffentlicht: (2025)
von: Cai, Chang, et al.
Veröffentlicht: (2025)
Interactive Text-to-Image Retrieval with Large Language Models: A Plug-and-Play Approach
von: Lee, Saehyung, et al.
Veröffentlicht: (2024)
von: Lee, Saehyung, et al.
Veröffentlicht: (2024)
Improving the Robustness of Distantly-Supervised Named Entity Recognition via Uncertainty-Aware Teacher Learning and Student-Student Collaborative Learning
von: Si, Shuzheng, et al.
Veröffentlicht: (2023)
von: Si, Shuzheng, et al.
Veröffentlicht: (2023)
Accelerating Inference in Large Language Models with a Unified Layer Skipping Strategy
von: Liu, Yijin, et al.
Veröffentlicht: (2024)
von: Liu, Yijin, et al.
Veröffentlicht: (2024)
Text-to-Image Rectified Flow as Plug-and-Play Priors
von: Yang, Xiaofeng, et al.
Veröffentlicht: (2024)
von: Yang, Xiaofeng, et al.
Veröffentlicht: (2024)
Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction
von: Zhao, Shiyu, et al.
Veröffentlicht: (2024)
von: Zhao, Shiyu, et al.
Veröffentlicht: (2024)
S2MDF: A Plug-And-Play Layer for Intersection-Free Multi-Object Signed Distance Fields
von: Mercadier, Deniz Sayin, et al.
Veröffentlicht: (2026)
von: Mercadier, Deniz Sayin, et al.
Veröffentlicht: (2026)
Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey
von: Chen, Liang, et al.
Veröffentlicht: (2024)
von: Chen, Liang, et al.
Veröffentlicht: (2024)
Predicting Rewards Alongside Tokens: Non-disruptive Parameter Insertion for Efficient Inference Intervention in Large Language Model
von: Yuan, Chenhan, et al.
Veröffentlicht: (2024)
von: Yuan, Chenhan, et al.
Veröffentlicht: (2024)
FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models
von: Li, Senmao, et al.
Veröffentlicht: (2025)
von: Li, Senmao, et al.
Veröffentlicht: (2025)
Self-Checker: Plug-and-Play Modules for Fact-Checking with Large Language Models
von: Li, Miaoran, et al.
Veröffentlicht: (2023)
von: Li, Miaoran, et al.
Veröffentlicht: (2023)
Enhanced Motion Forecasting with Plug-and-Play Multimodal Large Language Models
von: Luo, Katie, et al.
Veröffentlicht: (2025)
von: Luo, Katie, et al.
Veröffentlicht: (2025)
Diffusion Sampling Path Tells More: An Efficient Plug-and-Play Strategy for Sample Filtering
von: Wang, Sixian, et al.
Veröffentlicht: (2025)
von: Wang, Sixian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Spark of Vision-Language Intelligence: 2-Dimensional Autoregressive Transformer for Efficient Finegrained Image Generation
von: Chen, Liang, et al.
Veröffentlicht: (2024) -
Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think
von: Chen, Liang, et al.
Veröffentlicht: (2025) -
Looking Beyond Text: Reducing Language bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance
von: Zhao, Haozhe, et al.
Veröffentlicht: (2024) -
VLCache: Computing 2% Vision Tokens and Reusing 98% for Vision-Language Inference
von: Qin, Shengling, et al.
Veröffentlicht: (2025) -
Turbo: Informativity-Driven Acceleration Plug-In for Vision-Language Large Models
von: Ju, Chen, et al.
Veröffentlicht: (2024)