LightVLM: Acceleraing Large Multimodal Models with Pyramid Token Merging and KV Cache Compression
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Lianyu, Shang, Fanhua, Feng, Wei, Wan, Liang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
iLLaVA: An Image is Worth Fewer Than 1/3 Input Tokens in Large Multimodal Models
by: Hu, Lianyu, et al.
Published: (2024)
by: Hu, Lianyu, et al.
Published: (2024)
Deep Correlated Prompting for Visual Recognition with Missing Modalities
by: Hu, Lianyu, et al.
Published: (2024)
by: Hu, Lianyu, et al.
Published: (2024)
Pose-Guided Fine-Grained Sign Language Video Generation
by: Shi, Tongkai, et al.
Published: (2024)
by: Shi, Tongkai, et al.
Published: (2024)
Pyramid Forcing: Head-Aware Pyramid KV Cache Policy for High-Quality Long Video Generation
by: Chen, Jiayu, et al.
Published: (2026)
by: Chen, Jiayu, et al.
Published: (2026)
Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models
by: Ji, Yicheng, et al.
Published: (2026)
by: Ji, Yicheng, et al.
Published: (2026)
Completed Feature Disentanglement Learning for Multimodal MRIs Analysis
by: Liu, Tianling, et al.
Published: (2024)
by: Liu, Tianling, et al.
Published: (2024)
A Survey of Token Compression for Efficient Multimodal Large Language Models
by: Shao, Kele, et al.
Published: (2025)
by: Shao, Kele, et al.
Published: (2025)
CalibQuant: 1-Bit KV Cache Quantization for Multimodal LLMs
by: Han, Insu, et al.
Published: (2025)
by: Han, Insu, et al.
Published: (2025)
CorrNet+: Sign Language Recognition and Translation via Spatial-Temporal Correlation
by: Hu, Lianyu, et al.
Published: (2024)
by: Hu, Lianyu, et al.
Published: (2024)
CFCML: A Coarse-to-Fine Crossmodal Learning Framework For Disease Diagnosis Using Multimodal Images and Tabular Data
by: Liu, Tianling, et al.
Published: (2026)
by: Liu, Tianling, et al.
Published: (2026)
Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language Models
by: Liu, Xuyang, et al.
Published: (2025)
by: Liu, Xuyang, et al.
Published: (2025)
Head-Aware KV Cache Compression for Efficient Visual Autoregressive Modeling
by: Qin, Ziran, et al.
Published: (2025)
by: Qin, Ziran, et al.
Published: (2025)
Beyond Background Shift: Rethinking Instance Replay in Continual Semantic Segmentation
by: Yin, Hongmei, et al.
Published: (2025)
by: Yin, Hongmei, et al.
Published: (2025)
FlashVLM: Text-Guided Visual Token Selection for Large Multimodal Models
by: Cai, Kaitong, et al.
Published: (2025)
by: Cai, Kaitong, et al.
Published: (2025)
Can Visual Input Be Compressed? A Visual Token Compression Benchmark for Large Multimodal Models
by: Peng, Tianfan, et al.
Published: (2025)
by: Peng, Tianfan, et al.
Published: (2025)
Cached Adaptive Token Merging: Dynamic Token Reduction and Redundant Computation Elimination in Diffusion Model
by: Saghatchian, Omid, et al.
Published: (2025)
by: Saghatchian, Omid, et al.
Published: (2025)
PDD: Manifold-Prior Diverse Distillation for Medical Anomaly Detection
by: Lu, Xijun, et al.
Published: (2026)
by: Lu, Xijun, et al.
Published: (2026)
StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and Compression
by: Chen, Yilong, et al.
Published: (2025)
by: Chen, Yilong, et al.
Published: (2025)
LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval
by: Ning, Zhenyu, et al.
Published: (2025)
by: Ning, Zhenyu, et al.
Published: (2025)
LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models
by: Shang, Yuzhang, et al.
Published: (2024)
by: Shang, Yuzhang, et al.
Published: (2024)
AirCache: Activating Inter-modal Relevancy KV Cache Compression for Efficient Large Vision-Language Model Inference
by: Huang, Kai, et al.
Published: (2025)
by: Huang, Kai, et al.
Published: (2025)
PPE: Positional Preservation Embedding for Token Compression in Multimodal Large Language Models
by: Huang, Mouxiao, et al.
Published: (2025)
by: Huang, Mouxiao, et al.
Published: (2025)
LagKV: Lag-Relative Information of the KV Cache Tells Which Tokens Are Important
by: Liang, Manlai, et al.
Published: (2025)
by: Liang, Manlai, et al.
Published: (2025)
HybridToken-VLM: Hybrid Token Compression for Vision-Language Models
by: Zhang, Jusheng, et al.
Published: (2025)
by: Zhang, Jusheng, et al.
Published: (2025)
Revisiting Multimodal KV Cache Compression: A Frequency-Domain-Guided Outlier-KV-Aware Approach
by: Yang, Yaoxin, et al.
Published: (2025)
by: Yang, Yaoxin, et al.
Published: (2025)
PackCache: A Training-Free Acceleration Method for Unified Autoregressive Video Generation via Compact KV-Cache
by: Li, Kunyang, et al.
Published: (2026)
by: Li, Kunyang, et al.
Published: (2026)
InfoMerge: Information-aware Token Compression for Efficient Video Large Language Models
by: Liu, Xinxin, et al.
Published: (2026)
by: Liu, Xinxin, et al.
Published: (2026)
TokenCarve: Information-Preserving Visual Token Compression in Multimodal Large Language Models
by: Tan, Xudong, et al.
Published: (2025)
by: Tan, Xudong, et al.
Published: (2025)
Pyramid Token Pruning for High-Resolution Large Vision-Language Models via Region, Token, and Instruction-Guided Importance
by: Liang, Yuxuan, et al.
Published: (2025)
by: Liang, Yuxuan, et al.
Published: (2025)
LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token Merging
by: Shu, Zhijian, et al.
Published: (2025)
by: Shu, Zhijian, et al.
Published: (2025)
XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression
by: Su, Zunhai, et al.
Published: (2026)
by: Su, Zunhai, et al.
Published: (2026)
XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression
by: Su, Zunhai, et al.
Published: (2026)
by: Su, Zunhai, et al.
Published: (2026)
OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference
by: Chen, Wei, et al.
Published: (2024)
by: Chen, Wei, et al.
Published: (2024)
LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference
by: Wan, Zhongwei, et al.
Published: (2024)
by: Wan, Zhongwei, et al.
Published: (2024)
Dynamic Spatial-Temporal Aggregation for Skeleton-Aware Sign Language Recognition
by: Hu, Lianyu, et al.
Published: (2024)
by: Hu, Lianyu, et al.
Published: (2024)
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
Improving Continuous Sign Language Recognition with Adapted Image Models
by: Hu, Lianyu, et al.
Published: (2024)
by: Hu, Lianyu, et al.
Published: (2024)
Dynamic Pyramid Network for Efficient Multimodal Large Language Model
by: Ai, Hao, et al.
Published: (2025)
by: Ai, Hao, et al.
Published: (2025)
STaR-KV: Spatio-Temporal Adaptive Re-weighting for KV Cache Compression in GUI Vision-Language Models
by: Han, Yuhang, et al.
Published: (2026)
by: Han, Yuhang, et al.
Published: (2026)
Decouple and Cache: KV Cache Construction for Streaming Video Understanding
by: Pang, Zhanzhong, et al.
Published: (2026)
by: Pang, Zhanzhong, et al.
Published: (2026)
Similar Items
-
iLLaVA: An Image is Worth Fewer Than 1/3 Input Tokens in Large Multimodal Models
by: Hu, Lianyu, et al.
Published: (2024) -
Deep Correlated Prompting for Visual Recognition with Missing Modalities
by: Hu, Lianyu, et al.
Published: (2024) -
Pose-Guided Fine-Grained Sign Language Video Generation
by: Shi, Tongkai, et al.
Published: (2024) -
Pyramid Forcing: Head-Aware Pyramid KV Cache Policy for High-Quality Long Video Generation
by: Chen, Jiayu, et al.
Published: (2026) -
Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models
by: Ji, Yicheng, et al.
Published: (2026)