From Long Videos to Engaging Clips: A Human-Inspired Video Editing Framework with Multimodal Narrative Understanding
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wang, Xiangfeng, Li, Xiao, Wei, Yadong, Song, Xueyu, Song, Yang, Xia, Xiaoqiang, Zeng, Fangrui, Chen, Zaiyi, Liu, Liu, Xu, Gu, Xu, Tong |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
In-Context Learning with Unpaired Clips for Instruction-based Video Editing
par: Liao, Xinyao, et autres
Publié: (2025)
par: Liao, Xinyao, et autres
Publié: (2025)
Vidi: Large Multimodal Models for Video Understanding and Editing
par: Vidi Team, et autres
Publié: (2025)
par: Vidi Team, et autres
Publié: (2025)
VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges
par: Wang, Yuxuan, et autres
Publié: (2024)
par: Wang, Yuxuan, et autres
Publié: (2024)
Beyond Raw Videos: Understanding Edited Videos with Large Multimodal Model
par: Xu, Lu, et autres
Publié: (2024)
par: Xu, Lu, et autres
Publié: (2024)
Symphony: A Cognitively-Inspired Multi-Agent System for Long-Video Understanding
par: Yan, Haiyang, et autres
Publié: (2026)
par: Yan, Haiyang, et autres
Publié: (2026)
Think-Clip-Sample: Slow-Fast Frame Selection for Video Understanding
par: Tan, Wenhui, et autres
Publié: (2026)
par: Tan, Wenhui, et autres
Publié: (2026)
From Frames to Clips: Training-free Adaptive Key Clip Selection for Long-Form Video Understanding
par: Sun, Guangyu, et autres
Publié: (2025)
par: Sun, Guangyu, et autres
Publié: (2025)
Consistent Video Editing as Flow-Driven Image-to-Video Generation
par: Wang, Ge, et autres
Publié: (2025)
par: Wang, Ge, et autres
Publié: (2025)
In-Context Former: Lightning-fast Compressing Context for Large Language Model
par: Wang, Xiangfeng, et autres
Publié: (2024)
par: Wang, Xiangfeng, et autres
Publié: (2024)
Towards Effective Long Video Understanding of Multimodal Large Language Models via One-shot Clip Retrieval
par: Chen, Tao, et autres
Publié: (2025)
par: Chen, Tao, et autres
Publié: (2025)
STORM: Token-Efficient Long Video Understanding for Multimodal LLMs
par: Jiang, Jindong, et autres
Publié: (2025)
par: Jiang, Jindong, et autres
Publié: (2025)
Minimal Clips, Maximum Salience: Long Video Summarization via Key Moment Extraction
par: Pennec, Galann, et autres
Publié: (2025)
par: Pennec, Galann, et autres
Publié: (2025)
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization
par: Jin, Yang, et autres
Publié: (2024)
par: Jin, Yang, et autres
Publié: (2024)
Tele-Omni: a Unified Multimodal Framework for Video Generation and Editing
par: Liu, Jialun, et autres
Publié: (2026)
par: Liu, Jialun, et autres
Publié: (2026)
UniVideo: Unified Understanding, Generation, and Editing for Videos
par: Wei, Cong, et autres
Publié: (2025)
par: Wei, Cong, et autres
Publié: (2025)
Knowledge Editing for Large Language Models: A Survey
par: Wang, Song, et autres
Publié: (2023)
par: Wang, Song, et autres
Publié: (2023)
VideoSEAL: Mitigating Evidence Misalignment in Agentic Long Video Understanding by Decoupling Answer Authority
par: Qiu, Chenhao, et autres
Publié: (2026)
par: Qiu, Chenhao, et autres
Publié: (2026)
VideoExplorer: Think With Videos For Agentic Long-Video Understanding
par: Yuan, Huaying, et autres
Publié: (2025)
par: Yuan, Huaying, et autres
Publié: (2025)
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders
par: Fang, Bo, et autres
Publié: (2025)
par: Fang, Bo, et autres
Publié: (2025)
TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
par: Cai, Mu, et autres
Publié: (2024)
par: Cai, Mu, et autres
Publié: (2024)
1st Place Winner of the 2024 Pixel-level Video Understanding in the Wild (CVPR'24 PVUW) Challenge in Video Panoptic Segmentation and Best Long Video Consistency of Video Semantic Segmentation
par: Liu, Qingfeng, et autres
Publié: (2024)
par: Liu, Qingfeng, et autres
Publié: (2024)
REGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video Editing
par: Xu, Weihan, et autres
Publié: (2025)
par: Xu, Weihan, et autres
Publié: (2025)
Jailbreaking Multimodal Large Language Models using Multi-Clip Video
par: Kang, Choongwon, et autres
Publié: (2026)
par: Kang, Choongwon, et autres
Publié: (2026)
Divide-then-Diagnose: Weaving Clinician-Inspired Contexts for Ultra-Long Capsule Endoscopy Videos
par: Liu, Bowen, et autres
Publié: (2026)
par: Liu, Bowen, et autres
Publié: (2026)
Towards Human-like Multimodal Conversational Agent by Generating Engaging Speech
par: Kim, Taesoo, et autres
Publié: (2025)
par: Kim, Taesoo, et autres
Publié: (2025)
OmChat: A Recipe to Train Multimodal Language Models with Strong Long Context and Video Understanding
par: Zhao, Tiancheng, et autres
Publié: (2024)
par: Zhao, Tiancheng, et autres
Publié: (2024)
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding
par: Xu, Weili, et autres
Publié: (2025)
par: Xu, Weili, et autres
Publié: (2025)
Unleashing Hour-Scale Video Training for Long Video-Language Understanding
par: Lin, Jingyang, et autres
Publié: (2025)
par: Lin, Jingyang, et autres
Publié: (2025)
Learning Compact Video Representations for Efficient Long-form Video Understanding in Large Multimodal Models
par: Chen, Yuxiao, et autres
Publié: (2026)
par: Chen, Yuxiao, et autres
Publié: (2026)
RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives
par: Yoon, Jaehong, et autres
Publié: (2024)
par: Yoon, Jaehong, et autres
Publié: (2024)
Understanding Long Videos with Multimodal Language Models
par: Ranasinghe, Kanchana, et autres
Publié: (2024)
par: Ranasinghe, Kanchana, et autres
Publié: (2024)
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
par: Ren, Shuhuai, et autres
Publié: (2023)
par: Ren, Shuhuai, et autres
Publié: (2023)
Long Video Understanding with Learnable Retrieval in Video-Language Models
par: Xu, Jiaqi, et autres
Publié: (2023)
par: Xu, Jiaqi, et autres
Publié: (2023)
VideoAuteur: Towards Long Narrative Video Generation
par: Xiao, Junfei, et autres
Publié: (2025)
par: Xiao, Junfei, et autres
Publié: (2025)
Apollo: An Exploration of Video Understanding in Large Multimodal Models
par: Zohar, Orr, et autres
Publié: (2024)
par: Zohar, Orr, et autres
Publié: (2024)
Shot Sequence Ordering for Video Editing: Benchmarks, Metrics, and Cinematology-Inspired Computing Methods
par: Li, Yuzhi, et autres
Publié: (2025)
par: Li, Yuzhi, et autres
Publié: (2025)
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
par: Chen, Guo, et autres
Publié: (2024)
par: Chen, Guo, et autres
Publié: (2024)
RGNet: A Unified Clip Retrieval and Grounding Network for Long Videos
par: Hannan, Tanveer, et autres
Publié: (2023)
par: Hannan, Tanveer, et autres
Publié: (2023)
Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding
par: Wang, Youze, et autres
Publié: (2025)
par: Wang, Youze, et autres
Publié: (2025)
Towards Long Video Understanding via Fine-detailed Video Story Generation
par: You, Zeng, et autres
Publié: (2024)
par: You, Zeng, et autres
Publié: (2024)
Documents similaires
-
In-Context Learning with Unpaired Clips for Instruction-based Video Editing
par: Liao, Xinyao, et autres
Publié: (2025) -
Vidi: Large Multimodal Models for Video Understanding and Editing
par: Vidi Team, et autres
Publié: (2025) -
VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges
par: Wang, Yuxuan, et autres
Publié: (2024) -
Beyond Raw Videos: Understanding Edited Videos with Large Multimodal Model
par: Xu, Lu, et autres
Publié: (2024) -
Symphony: A Cognitively-Inspired Multi-Agent System for Long-Video Understanding
par: Yan, Haiyang, et autres
Publié: (2026)