Beyond Raw Videos: Understanding Edited Videos with Large Multimodal Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Lu, Zhu, Sijie, Li, Chunyuan, Kuo, Chia-Wen, Chen, Fan, Wang, Xinyao, Chen, Guang, Du, Dawei, Yuan, Ye, Wen, Longyin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing
von: Deng, Andong, et al.
Veröffentlicht: (2026)
von: Deng, Andong, et al.
Veröffentlicht: (2026)
Vidi: Large Multimodal Models for Video Understanding and Editing
von: Vidi Team, et al.
Veröffentlicht: (2025)
von: Vidi Team, et al.
Veröffentlicht: (2025)
D-Attn: Decomposed Attention for Large Vision-and-Language Models
von: Kuo, Chia-Wen, et al.
Veröffentlicht: (2025)
von: Kuo, Chia-Wen, et al.
Veröffentlicht: (2025)
CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts
von: Li, Jiachen, et al.
Veröffentlicht: (2024)
von: Li, Jiachen, et al.
Veröffentlicht: (2024)
Edit3K: Universal Representation Learning for Video Editing Components
von: Gu, Xin, et al.
Veröffentlicht: (2024)
von: Gu, Xin, et al.
Veröffentlicht: (2024)
Vidi2.5: Large Multimodal Models for Video Understanding and Creation
von: Vidi Team, et al.
Veröffentlicht: (2025)
von: Vidi Team, et al.
Veröffentlicht: (2025)
Where do Large Vision-Language Models Look at when Answering Questions?
von: Xing, Xiaoying, et al.
Veröffentlicht: (2025)
von: Xing, Xiaoying, et al.
Veröffentlicht: (2025)
Thinking With Bounding Boxes: Enhancing Spatio-Temporal Video Grounding via Reinforcement Fine-Tuning
von: Gu, Xin, et al.
Veröffentlicht: (2025)
von: Gu, Xin, et al.
Veröffentlicht: (2025)
SuperEdit: Rectifying and Facilitating Supervision for Instruction-Based Image Editing
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
Multi-Reward as Condition for Instruction-based Image Editing
von: Gu, Xin, et al.
Veröffentlicht: (2024)
von: Gu, Xin, et al.
Veröffentlicht: (2024)
CyberV: Cybernetics for Test-time Scaling in Video Understanding
von: Meng, Jiahao, et al.
Veröffentlicht: (2025)
von: Meng, Jiahao, et al.
Veröffentlicht: (2025)
Accurate and Fast Compressed Video Captioning
von: Shen, Yaojie, et al.
Veröffentlicht: (2023)
von: Shen, Yaojie, et al.
Veröffentlicht: (2023)
UniVideo: Unified Understanding, Generation, and Editing for Videos
von: Wei, Cong, et al.
Veröffentlicht: (2025)
von: Wei, Cong, et al.
Veröffentlicht: (2025)
Personalized Video Summarization by Multimodal Video Understanding
von: Chen, Brian, et al.
Veröffentlicht: (2024)
von: Chen, Brian, et al.
Veröffentlicht: (2024)
Engagement Prediction of Short Videos with Large Multimodal Models
von: Sun, Wei, et al.
Veröffentlicht: (2025)
von: Sun, Wei, et al.
Veröffentlicht: (2025)
Referring Layer Decomposition
von: Chen, Fangyi, et al.
Veröffentlicht: (2026)
von: Chen, Fangyi, et al.
Veröffentlicht: (2026)
From Long Videos to Engaging Clips: A Human-Inspired Video Editing Framework with Multimodal Narrative Understanding
von: Wang, Xiangfeng, et al.
Veröffentlicht: (2025)
von: Wang, Xiangfeng, et al.
Veröffentlicht: (2025)
Learning Compact Video Representations for Efficient Long-form Video Understanding in Large Multimodal Models
von: Chen, Yuxiao, et al.
Veröffentlicht: (2026)
von: Chen, Yuxiao, et al.
Veröffentlicht: (2026)
A Survey on Video Temporal Grounding with Multimodal Large Language Model
von: Wu, Jianlong, et al.
Veröffentlicht: (2025)
von: Wu, Jianlong, et al.
Veröffentlicht: (2025)
VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding
von: Fan, Yue, et al.
Veröffentlicht: (2024)
von: Fan, Yue, et al.
Veröffentlicht: (2024)
SAMA: Factorized Semantic Anchoring and Motion Alignment for Instruction-Guided Video Editing
von: Zhang, Xinyao, et al.
Veröffentlicht: (2026)
von: Zhang, Xinyao, et al.
Veröffentlicht: (2026)
TGT: Text-Grounded Trajectories for Locally Controlled Video Generation
von: Zhang, Guofeng, et al.
Veröffentlicht: (2025)
von: Zhang, Guofeng, et al.
Veröffentlicht: (2025)
Beyond Rigid: Benchmarking Non-Rigid Video Editing
von: Qu, Bingzheng, et al.
Veröffentlicht: (2026)
von: Qu, Bingzheng, et al.
Veröffentlicht: (2026)
VEU-Bench: Towards Comprehensive Understanding of Video Editing
von: Li, Bozheng, et al.
Veröffentlicht: (2025)
von: Li, Bozheng, et al.
Veröffentlicht: (2025)
LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models
von: Wei, Hongchen, et al.
Veröffentlicht: (2025)
von: Wei, Hongchen, et al.
Veröffentlicht: (2025)
REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding
von: Li, Jiaze, et al.
Veröffentlicht: (2025)
von: Li, Jiaze, et al.
Veröffentlicht: (2025)
Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding
von: Wang, Youze, et al.
Veröffentlicht: (2025)
von: Wang, Youze, et al.
Veröffentlicht: (2025)
Alignment-free Raw Video Demoireing
von: Xu, Shuning, et al.
Veröffentlicht: (2024)
von: Xu, Shuning, et al.
Veröffentlicht: (2024)
REGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video Editing
von: Xu, Weihan, et al.
Veröffentlicht: (2025)
von: Xu, Weihan, et al.
Veröffentlicht: (2025)
Structured Context Learning for Generic Event Boundary Detection
von: Gu, Xin, et al.
Veröffentlicht: (2025)
von: Gu, Xin, et al.
Veröffentlicht: (2025)
EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation
von: Li, Bingxuan, et al.
Veröffentlicht: (2025)
von: Li, Bingxuan, et al.
Veröffentlicht: (2025)
Long Video Understanding with Learnable Retrieval in Video-Language Models
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
Mixup Helps Understanding Multimodal Video Better
von: Ma, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Ma, Xiaoyu, et al.
Veröffentlicht: (2025)
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
von: Zhang, Boqiang, et al.
Veröffentlicht: (2025)
von: Zhang, Boqiang, et al.
Veröffentlicht: (2025)
Fewer Tokens and Fewer Videos: Extending Video Understanding Abilities in Large Vision-Language Models
von: Chen, Shimin, et al.
Veröffentlicht: (2024)
von: Chen, Shimin, et al.
Veröffentlicht: (2024)
Towards Sparse Video Understanding and Reasoning
von: Xu, Chenwei, et al.
Veröffentlicht: (2026)
von: Xu, Chenwei, et al.
Veröffentlicht: (2026)
VideoGrain: Modulating Space-Time Attention for Multi-grained Video Editing
von: Yang, Xiangpeng, et al.
Veröffentlicht: (2025)
von: Yang, Xiangpeng, et al.
Veröffentlicht: (2025)
In-Context Learning with Unpaired Clips for Instruction-based Video Editing
von: Liao, Xinyao, et al.
Veröffentlicht: (2025)
von: Liao, Xinyao, et al.
Veröffentlicht: (2025)
VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding
von: Zhang, Zhihong, et al.
Veröffentlicht: (2025)
von: Zhang, Zhihong, et al.
Veröffentlicht: (2025)
Beyond Single-Sample: Reliable Multi-Sample Distillation for Video Understanding
von: Li, Songlin, et al.
Veröffentlicht: (2026)
von: Li, Songlin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing
von: Deng, Andong, et al.
Veröffentlicht: (2026) -
Vidi: Large Multimodal Models for Video Understanding and Editing
von: Vidi Team, et al.
Veröffentlicht: (2025) -
D-Attn: Decomposed Attention for Large Vision-and-Language Models
von: Kuo, Chia-Wen, et al.
Veröffentlicht: (2025) -
CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts
von: Li, Jiachen, et al.
Veröffentlicht: (2024) -
Edit3K: Universal Representation Learning for Video Editing Components
von: Gu, Xin, et al.
Veröffentlicht: (2024)