Progressive Video Condensation with MLLM Agent for Long-form Video Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yin, Yufei, Xing, Yuchen, Meng, Qianke, Chen, Minghao, Yang, Yan, Yu, Zhou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding
von: Yin, Yufei, et al.
Veröffentlicht: (2025)
von: Yin, Yufei, et al.
Veröffentlicht: (2025)
LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents
von: Chen, Boyu, et al.
Veröffentlicht: (2025)
von: Chen, Boyu, et al.
Veröffentlicht: (2025)
MSJoE: Jointly Evolving MLLM and Sampler for Efficient Long-Form Video Understanding
von: Tan, Wenhui, et al.
Veröffentlicht: (2026)
von: Tan, Wenhui, et al.
Veröffentlicht: (2026)
VCA: Video Curious Agent for Long Video Understanding
von: Yang, Zeyuan, et al.
Veröffentlicht: (2024)
von: Yang, Zeyuan, et al.
Veröffentlicht: (2024)
Multiple Consistent 2D-3D Mappings for Robust Zero-Shot 3D Visual Grounding
von: Yin, Yufei, et al.
Veröffentlicht: (2026)
von: Yin, Yufei, et al.
Veröffentlicht: (2026)
Video Token Merging for Long-form Video Understanding
von: Lee, Seon-Ho, et al.
Veröffentlicht: (2024)
von: Lee, Seon-Ho, et al.
Veröffentlicht: (2024)
Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding
von: Shu, Yan, et al.
Veröffentlicht: (2024)
von: Shu, Yan, et al.
Veröffentlicht: (2024)
Learning Compact Video Representations for Efficient Long-form Video Understanding in Large Multimodal Models
von: Chen, Yuxiao, et al.
Veröffentlicht: (2026)
von: Chen, Yuxiao, et al.
Veröffentlicht: (2026)
TokensGen: Harnessing Condensed Tokens for Long Video Generation
von: Ouyang, Wenqi, et al.
Veröffentlicht: (2025)
von: Ouyang, Wenqi, et al.
Veröffentlicht: (2025)
EEA: Exploration-Exploitation Agent for Long Video Understanding
von: Yang, Te, et al.
Veröffentlicht: (2025)
von: Yang, Te, et al.
Veröffentlicht: (2025)
Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video Understanding
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
VideoAgent: Long-form Video Understanding with Large Language Model as Agent
von: Wang, Xiaohan, et al.
Veröffentlicht: (2024)
von: Wang, Xiaohan, et al.
Veröffentlicht: (2024)
Long Video Understanding with Learnable Retrieval in Video-Language Models
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
Memory-enhanced Retrieval Augmentation for Long Video Understanding
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
Test-Time Temporal Sampling for Efficient MLLM Video Understanding
von: Wang, Kaibin, et al.
Veröffentlicht: (2025)
von: Wang, Kaibin, et al.
Veröffentlicht: (2025)
SciEducator: Scientific Video Understanding and Educating via Deming-Cycle Multi-Agent System
von: Xu, Zhiyu, et al.
Veröffentlicht: (2025)
von: Xu, Zhiyu, et al.
Veröffentlicht: (2025)
Spacewalk-18: A Benchmark for Multimodal and Long-form Procedural Video Understanding in Novel Domains
von: Tang, Zitian, et al.
Veröffentlicht: (2023)
von: Tang, Zitian, et al.
Veröffentlicht: (2023)
VirtueBench: Evaluating Trustworthiness under Uncertainty in Long Video Understanding
von: Yu, Xueqing, et al.
Veröffentlicht: (2026)
von: Yu, Xueqing, et al.
Veröffentlicht: (2026)
REALM: An MLLM-Agent Framework for Open World 3D Reasoning Segmentation and Editing on Gaussian Splatting
von: Shi, Changyue, et al.
Veröffentlicht: (2025)
von: Shi, Changyue, et al.
Veröffentlicht: (2025)
VideoPro: Adaptive Program Reasoning for Long Video Understanding
von: Li, Chenglin, et al.
Veröffentlicht: (2025)
von: Li, Chenglin, et al.
Veröffentlicht: (2025)
TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding
von: Tang, Canhui, et al.
Veröffentlicht: (2025)
von: Tang, Canhui, et al.
Veröffentlicht: (2025)
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification
von: Qin, Minghao, et al.
Veröffentlicht: (2025)
von: Qin, Minghao, et al.
Veröffentlicht: (2025)
VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video Understanding
von: Waheed, Abdul, et al.
Veröffentlicht: (2025)
von: Waheed, Abdul, et al.
Veröffentlicht: (2025)
Efficient Motion-Aware Video MLLM
von: Zhao, Zijia, et al.
Veröffentlicht: (2025)
von: Zhao, Zijia, et al.
Veröffentlicht: (2025)
Omni-Video 2: Scaling MLLM-Conditioned Diffusion for Unified Video Generation and Editing
von: Yang, Hao, et al.
Veröffentlicht: (2026)
von: Yang, Hao, et al.
Veröffentlicht: (2026)
MLVU: Benchmarking Multi-task Long Video Understanding
von: Zhou, Junjie, et al.
Veröffentlicht: (2024)
von: Zhou, Junjie, et al.
Veröffentlicht: (2024)
Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow
von: Liu, Ruyang, et al.
Veröffentlicht: (2025)
von: Liu, Ruyang, et al.
Veröffentlicht: (2025)
AdaptToken: Entropy-based Adaptive Token Selection for MLLM Long Video Understanding
von: Qi, Haozhe, et al.
Veröffentlicht: (2026)
von: Qi, Haozhe, et al.
Veröffentlicht: (2026)
MR. Video: "MapReduce" is the Principle for Long Video Understanding
von: Pang, Ziqi, et al.
Veröffentlicht: (2025)
von: Pang, Ziqi, et al.
Veröffentlicht: (2025)
Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
MLLM as Video Narrator: Mitigating Modality Imbalance in Video Moment Retrieval
von: Cai, Weitong, et al.
Veröffentlicht: (2024)
von: Cai, Weitong, et al.
Veröffentlicht: (2024)
VideoAgent2: Enhancing the LLM-Based Agent System for Long-Form Video Understanding by Uncertainty-Aware CoT
von: Zhi, Zhuo, et al.
Veröffentlicht: (2025)
von: Zhi, Zhuo, et al.
Veröffentlicht: (2025)
Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding
von: Zhang, Xiaoyi, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaoyi, et al.
Veröffentlicht: (2025)
VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning
von: Gao, Zhe, et al.
Veröffentlicht: (2026)
von: Gao, Zhe, et al.
Veröffentlicht: (2026)
Task-Aware KV Compression For Cost-Effective Long Video Understanding
von: Qin, Minghao, et al.
Veröffentlicht: (2025)
von: Qin, Minghao, et al.
Veröffentlicht: (2025)
Symphony: A Cognitively-Inspired Multi-Agent System for Long-Video Understanding
von: Yan, Haiyang, et al.
Veröffentlicht: (2026)
von: Yan, Haiyang, et al.
Veröffentlicht: (2026)
In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting
von: Peng, Taiying, et al.
Veröffentlicht: (2025)
von: Peng, Taiying, et al.
Veröffentlicht: (2025)
VideoExplorer: Think With Videos For Agentic Long-Video Understanding
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
MASRA: MLLM-Assisted Semantic-Relational Consistent Alignment for Video Temporal Grounding
von: Ran, Ran, et al.
Veröffentlicht: (2026)
von: Ran, Ran, et al.
Veröffentlicht: (2026)
Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis
von: Yang, Qize, et al.
Veröffentlicht: (2025)
von: Yang, Qize, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding
von: Yin, Yufei, et al.
Veröffentlicht: (2025) -
LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents
von: Chen, Boyu, et al.
Veröffentlicht: (2025) -
MSJoE: Jointly Evolving MLLM and Sampler for Efficient Long-Form Video Understanding
von: Tan, Wenhui, et al.
Veröffentlicht: (2026) -
VCA: Video Curious Agent for Long Video Understanding
von: Yang, Zeyuan, et al.
Veröffentlicht: (2024) -
Multiple Consistent 2D-3D Mappings for Robust Zero-Shot 3D Visual Grounding
von: Yin, Yufei, et al.
Veröffentlicht: (2026)