From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zou, Heqing, Luo, Tianze, Xie, Guiyang, Victor, Zhang, Lv, Fengmao, Wang, Guangcong, Chen, Junyang, Wang, Zhuochen, Zhang, Hansheng, Zhang, Huaijian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding
von: Zou, Heqing, et al.
Veröffentlicht: (2025)
von: Zou, Heqing, et al.
Veröffentlicht: (2025)
Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning
von: Chen, Ruizhe, et al.
Veröffentlicht: (2025)
von: Chen, Ruizhe, et al.
Veröffentlicht: (2025)
Video-STR: Reinforcing MLLMs in Video Spatio-Temporal Reasoning with Relation Graph
von: Wang, Wentao, et al.
Veröffentlicht: (2025)
von: Wang, Wentao, et al.
Veröffentlicht: (2025)
Text-based Talking Video Editing with Cascaded Conditional Diffusion
von: Han, Bo, et al.
Veröffentlicht: (2024)
von: Han, Bo, et al.
Veröffentlicht: (2024)
Large Language Models Meet Contrastive Learning: Zero-Shot Emotion Recognition Across Languages
von: Zou, Heqing, et al.
Veröffentlicht: (2025)
von: Zou, Heqing, et al.
Veröffentlicht: (2025)
Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation
von: Chen, Yuheng, et al.
Veröffentlicht: (2026)
von: Chen, Yuheng, et al.
Veröffentlicht: (2026)
MultiModal Action Conditioned Video Generation
von: Li, Yichen, et al.
Veröffentlicht: (2025)
von: Li, Yichen, et al.
Veröffentlicht: (2025)
LongHalQA: Long-Context Hallucination Evaluation for MultiModal Large Language Models
von: Qiu, Han, et al.
Veröffentlicht: (2024)
von: Qiu, Han, et al.
Veröffentlicht: (2024)
M3: 3D-Spatial MultiModal Memory
von: Zou, Xueyan, et al.
Veröffentlicht: (2025)
von: Zou, Xueyan, et al.
Veröffentlicht: (2025)
Early Joint Learning of Emotion Information Makes MultiModal Model Understand You Better
von: Ge, Mengying, et al.
Veröffentlicht: (2024)
von: Ge, Mengying, et al.
Veröffentlicht: (2024)
ControlEdit: A MultiModal Local Clothing Image Editing Method
von: Cheng, Di, et al.
Veröffentlicht: (2024)
von: Cheng, Di, et al.
Veröffentlicht: (2024)
MMA-Diffusion: MultiModal Attack on Diffusion Models
von: Yang, Yijun, et al.
Veröffentlicht: (2023)
von: Yang, Yijun, et al.
Veröffentlicht: (2023)
DiffV2IR: Visible-to-Infrared Diffusion Model via Vision-Language Understanding
von: Ran, Lingyan, et al.
Veröffentlicht: (2025)
von: Ran, Lingyan, et al.
Veröffentlicht: (2025)
MLlm-DR: Towards Explainable Depression Recognition with MultiModal Large Language Models
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
MMCORE: MultiModal COnnection with Representation Aligned Latent Embeddings
von: Li, Zijie, et al.
Veröffentlicht: (2026)
von: Li, Zijie, et al.
Veröffentlicht: (2026)
Weakly-supervised Domain Adaption for Aspect Extraction via Multi-level Interaction Transfer
von: Liang, Tao, et al.
Veröffentlicht: (2020)
von: Liang, Tao, et al.
Veröffentlicht: (2020)
MM-LLMs: Recent Advances in MultiModal Large Language Models
von: Zhang, Duzhen, et al.
Veröffentlicht: (2024)
von: Zhang, Duzhen, et al.
Veröffentlicht: (2024)
LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
CapS-Adapter: Caption-based MultiModal Adapter in Zero-Shot Classification
von: Wang, Qijie, et al.
Veröffentlicht: (2024)
von: Wang, Qijie, et al.
Veröffentlicht: (2024)
On the Feasibility of Using MultiModal LLMs to Execute AR Social Engineering Attacks
von: Bi, Ting, et al.
Veröffentlicht: (2025)
von: Bi, Ting, et al.
Veröffentlicht: (2025)
Re-M3Dr: Rebalanced MultiModal Mean Deviation Regression
von: Yin, Haojie, et al.
Veröffentlicht: (2026)
von: Yin, Haojie, et al.
Veröffentlicht: (2026)
MultiModal Fine-tuning with Synthetic Captions
von: Enomoto, Shohei, et al.
Veröffentlicht: (2026)
von: Enomoto, Shohei, et al.
Veröffentlicht: (2026)
M$^2$CD: A Unified MultiModal Framework for Optical-SAR Change Detection with Mixture of Experts and Self-Distillation
von: Liu, Ziyuan, et al.
Veröffentlicht: (2025)
von: Liu, Ziyuan, et al.
Veröffentlicht: (2025)
MMViR: A Multi-Modal and Multi-Granularity Representation for Long-range Video Understanding
von: Li, Zizhong, et al.
Veröffentlicht: (2026)
von: Li, Zizhong, et al.
Veröffentlicht: (2026)
Piculet: Specialized Models-Guided Hallucination Decrease for MultiModal Large Language Models
von: Wang, Kohou, et al.
Veröffentlicht: (2024)
von: Wang, Kohou, et al.
Veröffentlicht: (2024)
MMRPT: MultiModal Reinforcement Pre-Training via Masked Vision-Dependent Reasoning
von: Zheng, Xuhui, et al.
Veröffentlicht: (2025)
von: Zheng, Xuhui, et al.
Veröffentlicht: (2025)
LVCHAT: Facilitating Long Video Comprehension
von: Wang, Yu, et al.
Veröffentlicht: (2024)
von: Wang, Yu, et al.
Veröffentlicht: (2024)
Frame-Level Captions for Long Video Generation with Complex Multi Scenes
von: Zheng, Guangcong, et al.
Veröffentlicht: (2025)
von: Zheng, Guangcong, et al.
Veröffentlicht: (2025)
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
von: Hao, Yunzhuo, et al.
Veröffentlicht: (2025)
von: Hao, Yunzhuo, et al.
Veröffentlicht: (2025)
Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought
von: Zhang, Shuyi, et al.
Veröffentlicht: (2025)
von: Zhang, Shuyi, et al.
Veröffentlicht: (2025)
Colorectal Polyp Segmentation in the Deep Learning Era: A Comprehensive Survey
von: Wu, Zhenyu, et al.
Veröffentlicht: (2024)
von: Wu, Zhenyu, et al.
Veröffentlicht: (2024)
VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding
von: He, Haichen, et al.
Veröffentlicht: (2026)
von: He, Haichen, et al.
Veröffentlicht: (2026)
Unleashing Hour-Scale Video Training for Long Video-Language Understanding
von: Lin, Jingyang, et al.
Veröffentlicht: (2025)
von: Lin, Jingyang, et al.
Veröffentlicht: (2025)
VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding
von: Zhang, Zhihong, et al.
Veröffentlicht: (2025)
von: Zhang, Zhihong, et al.
Veröffentlicht: (2025)
Visual-RolePlay: Universal Jailbreak Attack on MultiModal Large Language Models via Role-playing Image Character
von: Ma, Siyuan, et al.
Veröffentlicht: (2024)
von: Ma, Siyuan, et al.
Veröffentlicht: (2024)
An Empirical Study on Parameter-Efficient Fine-Tuning for MultiModal Large Language Models
von: Zhou, Xiongtao, et al.
Veröffentlicht: (2024)
von: Zhou, Xiongtao, et al.
Veröffentlicht: (2024)
Dementia Insights: A Context-Based MultiModal Approach
von: Mehdoui, Sahar Sinene, et al.
Veröffentlicht: (2025)
von: Mehdoui, Sahar Sinene, et al.
Veröffentlicht: (2025)
MM-MoralBench: A MultiModal Moral Evaluation Benchmark for Large Vision-Language Models
von: Yan, Bei, et al.
Veröffentlicht: (2024)
von: Yan, Bei, et al.
Veröffentlicht: (2024)
Achieving Hour‐Long Organic Afterglow in Donor‐Sensitizer‐Acceptor System via Molecular Engineering of Sensitizers
von: Yunlong Zou, et al.
Veröffentlicht: (2026)
von: Yunlong Zou, et al.
Veröffentlicht: (2026)
Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding
von: Chen, Wang, et al.
Veröffentlicht: (2026)
von: Chen, Wang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding
von: Zou, Heqing, et al.
Veröffentlicht: (2025) -
Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning
von: Chen, Ruizhe, et al.
Veröffentlicht: (2025) -
Video-STR: Reinforcing MLLMs in Video Spatio-Temporal Reasoning with Relation Graph
von: Wang, Wentao, et al.
Veröffentlicht: (2025) -
Text-based Talking Video Editing with Cascaded Conditional Diffusion
von: Han, Bo, et al.
Veröffentlicht: (2024) -
Large Language Models Meet Contrastive Learning: Zero-Shot Emotion Recognition Across Languages
von: Zou, Heqing, et al.
Veröffentlicht: (2025)