VidCompress: Memory-Enhanced Temporal Compression for Video Understanding in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Lan, Xiaohan, Yuan, Yitian, Jie, Zequn, Ma, Lin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability
by: Chen, Shimin, et al.
Published: (2024)
by: Chen, Shimin, et al.
Published: (2024)
SSNVC: Single Stream Neural Video Compression with Implicit Temporal Information
by: Wang, Feng, et al.
Published: (2024)
by: Wang, Feng, et al.
Published: (2024)
SMC++: Masked Learning of Unsupervised Video Semantic Compression
by: Tian, Yuan, et al.
Published: (2024)
by: Tian, Yuan, et al.
Published: (2024)
Spatio-Temporal Data Enhanced Vision-Language Model for Traffic Scene Understanding
by: Ma, Jingtian, et al.
Published: (2025)
by: Ma, Jingtian, et al.
Published: (2025)
Context Guided Transformer Entropy Modeling for Video Compression
by: Tong, Junlong, et al.
Published: (2025)
by: Tong, Junlong, et al.
Published: (2025)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
by: Qin, Bosheng, et al.
Published: (2023)
by: Qin, Bosheng, et al.
Published: (2023)
CounterVid: Counterfactual Video Generation for Mitigating Action and Temporal Hallucinations in Video-Language Models
by: Poppi, Tobia, et al.
Published: (2026)
by: Poppi, Tobia, et al.
Published: (2026)
Bridging Compressed Image Latents and Multimodal Large Language Models
by: Kao, Chia-Hao, et al.
Published: (2024)
by: Kao, Chia-Hao, et al.
Published: (2024)
Delving Deeper: Hierarchical Visual Perception for Robust Video-Text Retrieval
by: Xie, Zequn, et al.
Published: (2026)
by: Xie, Zequn, et al.
Published: (2026)
FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis
by: Liang, Feng, et al.
Published: (2023)
by: Liang, Feng, et al.
Published: (2023)
Memory-enhanced Retrieval Augmentation for Long Video Understanding
by: Yuan, Huaying, et al.
Published: (2025)
by: Yuan, Huaying, et al.
Published: (2025)
Adaptive Rate Control for Deep Video Compression with Rate-Distortion Prediction
by: Gu, Bowen, et al.
Published: (2024)
by: Gu, Bowen, et al.
Published: (2024)
ChatVTG: Video Temporal Grounding via Chat with Video Dialogue Large Language Models
by: Qu, Mengxue, et al.
Published: (2024)
by: Qu, Mengxue, et al.
Published: (2024)
RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving
by: Huang, Zhijian, et al.
Published: (2024)
by: Huang, Zhijian, et al.
Published: (2024)
Efficient and Generic Point Model for Lossless Point Cloud Attribute Compression
by: You, Kang, et al.
Published: (2024)
by: You, Kang, et al.
Published: (2024)
Rate-aware Compression for NeRF-based Volumetric Video
by: Zhang, Zhiyu, et al.
Published: (2024)
by: Zhang, Zhiyu, et al.
Published: (2024)
Hybrid Local-Global Context Learning for Neural Video Compression
by: Zhai, Yongqi, et al.
Published: (2024)
by: Zhai, Yongqi, et al.
Published: (2024)
A Tri-Dynamic Preprocessing Framework for UGC Video Compression
by: Zhao, Fei, et al.
Published: (2025)
by: Zhao, Fei, et al.
Published: (2025)
A Preprocessing Framework for Video Machine Vision under Compression
by: Zhao, Fei, et al.
Published: (2025)
by: Zhao, Fei, et al.
Published: (2025)
Context-Enhanced Video Moment Retrieval with Large Language Models
by: Liu, Weijia, et al.
Published: (2024)
by: Liu, Weijia, et al.
Published: (2024)
Enhancing 3D Gaussian Splatting Compression via Spatial Condition-based Prediction
by: Ma, Jingui, et al.
Published: (2025)
by: Ma, Jingui, et al.
Published: (2025)
Accelerated Event-Based Feature Detection and Compression for Surveillance Video Systems
by: Freeman, Andrew C., et al.
Published: (2023)
by: Freeman, Andrew C., et al.
Published: (2023)
Fewer Tokens and Fewer Videos: Extending Video Understanding Abilities in Large Vision-Language Models
by: Chen, Shimin, et al.
Published: (2024)
by: Chen, Shimin, et al.
Published: (2024)
A Hierarchical Compression Technique for 3D Gaussian Splatting Compression
by: Huang, He, et al.
Published: (2024)
by: Huang, He, et al.
Published: (2024)
GSVC: Efficient Video Representation and Compression Through 2D Gaussian Splatting
by: Wang, Longan, et al.
Published: (2025)
by: Wang, Longan, et al.
Published: (2025)
D-FCGS: Feedforward Compression of Dynamic Gaussian Splatting for Free-Viewpoint Videos
by: Zhang, Wenkang, et al.
Published: (2025)
by: Zhang, Wenkang, et al.
Published: (2025)
IsoSignVid2Aud: Sign Language Video to Audio Conversion without Text Intermediaries
by: Kavediya, Harsh, et al.
Published: (2025)
by: Kavediya, Harsh, et al.
Published: (2025)
GaussianForest: Hierarchical-Hybrid 3D Gaussian Splatting for Compressed Scene Modeling
by: Zhang, Fengyi, et al.
Published: (2024)
by: Zhang, Fengyi, et al.
Published: (2024)
Quantifying and Enhancing Multi-modal Robustness with Modality Preference
by: Yang, Zequn, et al.
Published: (2024)
by: Yang, Zequn, et al.
Published: (2024)
VidCtx: Context-aware Video Question Answering with Image Models
by: Goulas, Andreas, et al.
Published: (2024)
by: Goulas, Andreas, et al.
Published: (2024)
UniVid: Pyramid Diffusion Model for High Quality Video Generation
by: Xiao, Xinyu, et al.
Published: (2026)
by: Xiao, Xinyu, et al.
Published: (2026)
Compression of 3D Gaussian Splatting with Optimized Feature Planes and Standard Video Codecs
by: Lee, Soonbin, et al.
Published: (2025)
by: Lee, Soonbin, et al.
Published: (2025)
L-LBVC: Long-Term Motion Estimation and Prediction for Learned Bi-Directional Video Compression
by: Zhai, Yongqi, et al.
Published: (2025)
by: Zhai, Yongqi, et al.
Published: (2025)
MFQE 2.0: A New Approach for Multi-frame Quality Enhancement on Compressed Video
by: Xing, Qunliang, et al.
Published: (2019)
by: Xing, Qunliang, et al.
Published: (2019)
Disparity-based Stereo Image Compression with Aligned Cross-View Priors
by: Zhai, Yongqi, et al.
Published: (2022)
by: Zhai, Yongqi, et al.
Published: (2022)
STEAR: Layer-Aware Spatiotemporal Evidence Intervention for Hallucination Mitigation in Video Large Language Models
by: Fan, Linfeng, et al.
Published: (2026)
by: Fan, Linfeng, et al.
Published: (2026)
Towards Reproducible Learning-based Compression
by: Pang, Jiahao, et al.
Published: (2024)
by: Pang, Jiahao, et al.
Published: (2024)
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos
by: Kong, Fanheng, et al.
Published: (2025)
by: Kong, Fanheng, et al.
Published: (2025)
SPC-NeRF: Spatial Predictive Compression for Voxel Based Radiance Field
by: Song, Zetian, et al.
Published: (2024)
by: Song, Zetian, et al.
Published: (2024)
Predicting Satisfied User and Machine Ratio for Compressed Images: A Unified Approach
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
Similar Items
-
TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability
by: Chen, Shimin, et al.
Published: (2024) -
SSNVC: Single Stream Neural Video Compression with Implicit Temporal Information
by: Wang, Feng, et al.
Published: (2024) -
SMC++: Masked Learning of Unsupervised Video Semantic Compression
by: Tian, Yuan, et al.
Published: (2024) -
Spatio-Temporal Data Enhanced Vision-Language Model for Traffic Scene Understanding
by: Ma, Jingtian, et al.
Published: (2025) -
Context Guided Transformer Entropy Modeling for Video Compression
by: Tong, Junlong, et al.
Published: (2025)