Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Liu, Xuyang, Wang, Yiyu, Ma, Junpeng, Zhang, Linfeng |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models
par: Liu, Xuyang, et autres
Publié: (2025)
par: Liu, Xuyang, et autres
Publié: (2025)
Accelerating Streaming Video Large Language Models via Hierarchical Token Compression
par: Wang, Yiyu, et autres
Publié: (2025)
par: Wang, Yiyu, et autres
Publié: (2025)
V-CAST: Video Curvature-Aware Spatio-Temporal Pruning for Efficient Video Large Language Models
par: Lin, Xinying, et autres
Publié: (2026)
par: Lin, Xinying, et autres
Publié: (2026)
Plug-and-Play Versatile Compressed Video Enhancement
par: Zeng, Huimin, et autres
Publié: (2025)
par: Zeng, Huimin, et autres
Publié: (2025)
Prune2Drive: A Plug-and-Play Framework for Accelerating Vision-Language Models in Autonomous Driving
par: Xiong, Minhao, et autres
Publié: (2025)
par: Xiong, Minhao, et autres
Publié: (2025)
Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language Models
par: Liu, Xuyang, et autres
Publié: (2025)
par: Liu, Xuyang, et autres
Publié: (2025)
Plug-and-Play 1.x-Bit KV Cache Quantization for Video Large Language Models
par: Tao, Keda, et autres
Publié: (2025)
par: Tao, Keda, et autres
Publié: (2025)
Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generation
par: Zhu, Lunjie, et autres
Publié: (2026)
par: Zhu, Lunjie, et autres
Publié: (2026)
Bridging Visual Representation and Reinforcement Learning from Verifiable Rewards in Large Vision-Language Models
par: Han, Yuhang, et autres
Publié: (2026)
par: Han, Yuhang, et autres
Publié: (2026)
Stand-In: A Lightweight and Plug-and-Play Identity Control for Video Generation
par: Xue, Bowen, et autres
Publié: (2025)
par: Xue, Bowen, et autres
Publié: (2025)
ARM: A Learnable, Plug-and-Play Module for CLIP-based Open-vocabulary Semantic Segmentation
par: Liu, Ziquan, et autres
Publié: (2025)
par: Liu, Ziquan, et autres
Publié: (2025)
FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models
par: Li, Senmao, et autres
Publié: (2025)
par: Li, Senmao, et autres
Publié: (2025)
Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models
par: Chen, Jiaxing, et autres
Publié: (2024)
par: Chen, Jiaxing, et autres
Publié: (2024)
Generalizing Deepfake Video Detection with Plug-and-Play: Video-Level Blending and Spatiotemporal Adapter Tuning
par: Yan, Zhiyuan, et autres
Publié: (2024)
par: Yan, Zhiyuan, et autres
Publié: (2024)
An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
par: Chen, Liang, et autres
Publié: (2024)
par: Chen, Liang, et autres
Publié: (2024)
Variation-aware Vision Token Dropping for Faster Large Vision-Language Models
par: Chen, Junjie, et autres
Publié: (2025)
par: Chen, Junjie, et autres
Publié: (2025)
Enhanced Motion Forecasting with Plug-and-Play Multimodal Large Language Models
par: Luo, Katie, et autres
Publié: (2025)
par: Luo, Katie, et autres
Publié: (2025)
RAG-Adapter: A Plug-and-Play RAG-enhanced Framework for Long Video Understanding
par: Tan, Xichen, et autres
Publié: (2025)
par: Tan, Xichen, et autres
Publié: (2025)
IPCV: Information-Preserving Compression for MLLM Visual Encoders
par: Chen, Yuan, et autres
Publié: (2025)
par: Chen, Yuan, et autres
Publié: (2025)
VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control
par: Bian, Yuxuan, et autres
Publié: (2025)
par: Bian, Yuxuan, et autres
Publié: (2025)
Score-Based Turbo Message Passing for Plug-and-Play Compressive Imaging
par: Cai, Chang, et autres
Publié: (2025)
par: Cai, Chang, et autres
Publié: (2025)
VidCompress: Memory-Enhanced Temporal Compression for Video Understanding in Large Language Models
par: Lan, Xiaohan, et autres
Publié: (2024)
par: Lan, Xiaohan, et autres
Publié: (2024)
ReMA: A Training-Free Plug-and-Play Mixing Augmentation for Video Behavior Recognition
par: Cui, Feng-Qi, et autres
Publié: (2026)
par: Cui, Feng-Qi, et autres
Publié: (2026)
VideoCompressa: Data-Efficient Video Understanding via Joint Temporal Compression and Spatial Reconstruction
par: Wang, Shaobo, et autres
Publié: (2025)
par: Wang, Shaobo, et autres
Publié: (2025)
Turbo: Informativity-Driven Acceleration Plug-In for Vision-Language Large Models
par: Ju, Chen, et autres
Publié: (2024)
par: Ju, Chen, et autres
Publié: (2024)
Interactive Text-to-Image Retrieval with Large Language Models: A Plug-and-Play Approach
par: Lee, Saehyung, et autres
Publié: (2024)
par: Lee, Saehyung, et autres
Publié: (2024)
EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models
par: Yang, Yantai, et autres
Publié: (2025)
par: Yang, Yantai, et autres
Publié: (2025)
AccVideo: Accelerating Video Diffusion Model with Synthetic Dataset
par: Zhang, Haiyu, et autres
Publié: (2025)
par: Zhang, Haiyu, et autres
Publié: (2025)
Short-LVLM: Compressing and Accelerating Large Vision-Language Models by Pruning Redundant Layers
par: Ma, Ji, et autres
Publié: (2025)
par: Ma, Ji, et autres
Publié: (2025)
HIPPO: Accelerating Video Large Language Models Inference via Holistic-aware Parallel Speculative Decoding
par: Lv, Qitan, et autres
Publié: (2026)
par: Lv, Qitan, et autres
Publié: (2026)
PAT-VCM: Plug-and-Play Auxiliary Tokens for Video Coding for Machines
par: Jiang, Wei, et autres
Publié: (2026)
par: Jiang, Wei, et autres
Publié: (2026)
EvoStreaming: Your Offline Video Model Is a Natively Streaming Assistant
par: Wen, Zichen, et autres
Publié: (2026)
par: Wen, Zichen, et autres
Publié: (2026)
Inference Compute-Optimal Video Vision Language Models
par: Wang, Peiqi, et autres
Publié: (2025)
par: Wang, Peiqi, et autres
Publié: (2025)
SpeedUpNet: A Plug-and-Play Adapter Network for Accelerating Text-to-Image Diffusion Models
par: Chai, Weilong, et autres
Publié: (2023)
par: Chai, Weilong, et autres
Publié: (2023)
STEAR: Layer-Aware Spatiotemporal Evidence Intervention for Hallucination Mitigation in Video Large Language Models
par: Fan, Linfeng, et autres
Publié: (2026)
par: Fan, Linfeng, et autres
Publié: (2026)
InfoMerge: Information-aware Token Compression for Efficient Video Large Language Models
par: Liu, Xinxin, et autres
Publié: (2026)
par: Liu, Xinxin, et autres
Publié: (2026)
HiCache: A Plug-in Scaled-Hermite Upgrade for Taylor-Style Cache-then-Forecast Diffusion Acceleration
par: Feng, Liang, et autres
Publié: (2025)
par: Feng, Liang, et autres
Publié: (2025)
Seeing Clearly, Reasoning Confidently: Plug-and-Play Remedies for Vision Language Model Blindness
par: Hu, Xin, et autres
Publié: (2026)
par: Hu, Xin, et autres
Publié: (2026)
Plug-and-Play Diffusion Distillation
par: Hsiao, Yi-Ting, et autres
Publié: (2024)
par: Hsiao, Yi-Ting, et autres
Publié: (2024)
VideoLLM-online: Online Video Large Language Model for Streaming Video
par: Chen, Joya, et autres
Publié: (2024)
par: Chen, Joya, et autres
Publié: (2024)
Documents similaires
-
Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models
par: Liu, Xuyang, et autres
Publié: (2025) -
Accelerating Streaming Video Large Language Models via Hierarchical Token Compression
par: Wang, Yiyu, et autres
Publié: (2025) -
V-CAST: Video Curvature-Aware Spatio-Temporal Pruning for Efficient Video Large Language Models
par: Lin, Xinying, et autres
Publié: (2026) -
Plug-and-Play Versatile Compressed Video Enhancement
par: Zeng, Huimin, et autres
Publié: (2025) -
Prune2Drive: A Plug-and-Play Framework for Accelerating Vision-Language Models in Autonomous Driving
par: Xiong, Minhao, et autres
Publié: (2025)