EMCompress: Video-LLMs with Endomorphic Multimodal Compression
Fuente:
arXiv
Saved in:
| Main Authors: | Fan, Zheyu, Liu, Jiateng, Zhang, Yuji, Wang, Zihan, Fung, Yi R., Li, Manling, Ji, Heng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios
by: Shi, Yang, et al.
Published: (2025)
by: Shi, Yang, et al.
Published: (2025)
One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding
by: Zhang, Zheyu, et al.
Published: (2026)
by: Zhang, Zheyu, et al.
Published: (2026)
Towards Sparse Video Understanding and Reasoning
by: Xu, Chenwei, et al.
Published: (2026)
by: Xu, Chenwei, et al.
Published: (2026)
Accurate and Fast Compressed Video Captioning
by: Shen, Yaojie, et al.
Published: (2023)
by: Shen, Yaojie, et al.
Published: (2023)
Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding
by: Wang, Youze, et al.
Published: (2025)
by: Wang, Youze, et al.
Published: (2025)
Subjective and Objective Quality Assessment of Banding Artifacts on Compressed Videos
by: Zheng, Qi, et al.
Published: (2025)
by: Zheng, Qi, et al.
Published: (2025)
Generative Video Compression with One-Dimensional Latent Representation
by: Zheng, Zihan, et al.
Published: (2026)
by: Zheng, Zihan, et al.
Published: (2026)
MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
by: Peng, Tianhao, et al.
Published: (2025)
by: Peng, Tianhao, et al.
Published: (2025)
Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers
by: Su, Zhaochen, et al.
Published: (2025)
by: Su, Zhaochen, et al.
Published: (2025)
STORM: Token-Efficient Long Video Understanding for Multimodal LLMs
by: Jiang, Jindong, et al.
Published: (2025)
by: Jiang, Jindong, et al.
Published: (2025)
Generative Models at the Frontier of Compression: A Survey on Generative Face Video Coding
by: Chen, Bolin, et al.
Published: (2025)
by: Chen, Bolin, et al.
Published: (2025)
TempCompass: Do Video LLMs Really Understand Videos?
by: Liu, Yuanxin, et al.
Published: (2024)
by: Liu, Yuanxin, et al.
Published: (2024)
Bi-Directional Deep Contextual Video Compression
by: Sheng, Xihua, et al.
Published: (2024)
by: Sheng, Xihua, et al.
Published: (2024)
Unified Spatiotemporal Token Compression for Video-LLMs at Ultra-Low Retention
by: Du, Junhao, et al.
Published: (2026)
by: Du, Junhao, et al.
Published: (2026)
VisionSelector: End-to-End Learnable Visual Token Compression for Efficient Multimodal LLMs
by: Zhu, Jiaying, et al.
Published: (2025)
by: Zhu, Jiaying, et al.
Published: (2025)
T*: Re-thinking Temporal Search for Long-Form Video Understanding
by: Ye, Jinhui, et al.
Published: (2025)
by: Ye, Jinhui, et al.
Published: (2025)
Can Multimodal LLMs See Science Instruction? Benchmarking Pedagogical Reasoning in K-12 Classroom Videos
by: Shen, Yixuan, et al.
Published: (2026)
by: Shen, Yixuan, et al.
Published: (2026)
HallE-Control: Controlling Object Hallucination in Large Multimodal Models
by: Zhai, Bohan, et al.
Published: (2023)
by: Zhai, Bohan, et al.
Published: (2023)
PNVC: Towards Practical INR-based Video Compression
by: Gao, Ge, et al.
Published: (2024)
by: Gao, Ge, et al.
Published: (2024)
GFix: Perceptually Enhanced Gaussian Splatting Video Compression
by: Teng, Siyue, et al.
Published: (2025)
by: Teng, Siyue, et al.
Published: (2025)
Video Token Sparsification for Efficient Multimodal LLMs in Autonomous Driving
by: Ma, Yunsheng, et al.
Published: (2024)
by: Ma, Yunsheng, et al.
Published: (2024)
An LMM for Efficient Video Understanding via Reinforced Compression of Video Cubes
by: Qi, Ji, et al.
Published: (2025)
by: Qi, Ji, et al.
Published: (2025)
DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding
by: Zhang, Hongzhi, et al.
Published: (2025)
by: Zhang, Hongzhi, et al.
Published: (2025)
Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Models
by: Liu, Xuyang, et al.
Published: (2025)
by: Liu, Xuyang, et al.
Published: (2025)
Toward Cognitive Supersensing in Multimodal Large Language Model
by: Li, Boyi, et al.
Published: (2026)
by: Li, Boyi, et al.
Published: (2026)
UltraGen: High-Resolution Video Generation with Hierarchical Attention
by: Hu, Teng, et al.
Published: (2025)
by: Hu, Teng, et al.
Published: (2025)
Conditional Video Generation for High-Efficiency Video Compression
by: Yi, Fangqiu, et al.
Published: (2025)
by: Yi, Fangqiu, et al.
Published: (2025)
Flow-Guided Diffusion for Video Inpainting
by: Gu, Bohai, et al.
Published: (2023)
by: Gu, Bohai, et al.
Published: (2023)
Seeing Clearly without Training: Mitigating Hallucinations in Multimodal LLMs for Remote Sensing
by: Liu, Yi, et al.
Published: (2026)
by: Liu, Yi, et al.
Published: (2026)
ARGUS: Hallucination and Omission Evaluation in Video-LLMs
by: Rawal, Ruchit, et al.
Published: (2025)
by: Rawal, Ruchit, et al.
Published: (2025)
M3-CVC: Controllable Video Compression with Multimodal Generative Models
by: Wan, Rui, et al.
Published: (2024)
by: Wan, Rui, et al.
Published: (2024)
Recoverable Compression: A Multimodal Vision Token Recovery Mechanism Guided by Text Information
by: Chen, Yi, et al.
Published: (2024)
by: Chen, Yi, et al.
Published: (2024)
CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs
by: Zhou, Xingcheng, et al.
Published: (2026)
by: Zhou, Xingcheng, et al.
Published: (2026)
CREM: Compression-Driven Representation Enhancement for Multimodal Retrieval and Comprehension
by: Liu, Lihao, et al.
Published: (2026)
by: Liu, Lihao, et al.
Published: (2026)
Controllable Generative Video Compression
by: Ding, Ding, et al.
Published: (2026)
by: Ding, Ding, et al.
Published: (2026)
High-Fidelity Image Inpainting with Multimodal Guided GAN Inversion
by: Zhang, Libo, et al.
Published: (2025)
by: Zhang, Libo, et al.
Published: (2025)
Visually Descriptive Language Model for Vector Graphics Reasoning
by: Wang, Zhenhailong, et al.
Published: (2024)
by: Wang, Zhenhailong, et al.
Published: (2024)
OFA-Diffusion Compression: Compressing Diffusion Model in One-Shot Manner
by: Jiang, Haoyang, et al.
Published: (2026)
by: Jiang, Haoyang, et al.
Published: (2026)
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
by: Wang, Yi, et al.
Published: (2024)
by: Wang, Yi, et al.
Published: (2024)
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs
by: Cheng, Dabing, et al.
Published: (2025)
by: Cheng, Dabing, et al.
Published: (2025)
Similar Items
-
MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios
by: Shi, Yang, et al.
Published: (2025) -
One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding
by: Zhang, Zheyu, et al.
Published: (2026) -
Towards Sparse Video Understanding and Reasoning
by: Xu, Chenwei, et al.
Published: (2026) -
Accurate and Fast Compressed Video Captioning
by: Shen, Yaojie, et al.
Published: (2023) -
Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding
by: Wang, Youze, et al.
Published: (2025)