VideoDirector: Precise Video Editing via Text-to-Video Models
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yukun, Wang, Longguang, Ma, Zhiyuan, Hu, Qibin, Xu, Kai, Guo, Yulan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FlowDirector: Training-Free Flow Steering for Precise Text-to-Video Editing
by: Li, Guangzhao, et al.
Published: (2025)
by: Li, Guangzhao, et al.
Published: (2025)
MonoMobility: Zero-Shot 3D Mobility Analysis from Monocular Videos
by: Zhou, Hongyi, et al.
Published: (2025)
by: Zhou, Hongyi, et al.
Published: (2025)
EffiVED:Efficient Video Editing via Text-instruction Diffusion Models
by: Zhang, Zhenghao, et al.
Published: (2024)
by: Zhang, Zhenghao, et al.
Published: (2024)
Layout2Scene: 3D Semantic Layout Guided Scene Generation via Geometry and Appearance Diffusion Priors
by: Chen, Minglin, et al.
Published: (2025)
by: Chen, Minglin, et al.
Published: (2025)
FastVideoEdit: Leveraging Consistency Models for Efficient Text-to-Video Editing
by: Zhang, Youyuan, et al.
Published: (2024)
by: Zhang, Youyuan, et al.
Published: (2024)
InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision
by: Wang, Chenting, et al.
Published: (2025)
by: Wang, Chenting, et al.
Published: (2025)
Shape-for-Motion: Precise and Consistent Video Editing with 3D Proxy
by: Liu, Yuhao, et al.
Published: (2025)
by: Liu, Yuhao, et al.
Published: (2025)
CCEdit: Creative and Controllable Video Editing via Diffusion Models
by: Feng, Ruoyu, et al.
Published: (2023)
by: Feng, Ruoyu, et al.
Published: (2023)
Beyond Raw Videos: Understanding Edited Videos with Large Multimodal Model
by: Xu, Lu, et al.
Published: (2024)
by: Xu, Lu, et al.
Published: (2024)
VideoMage: Multi-Subject and Motion Customization of Text-to-Video Diffusion Models
by: Huang, Chi-Pin, et al.
Published: (2025)
by: Huang, Chi-Pin, et al.
Published: (2025)
VideoCoF: Unified Video Editing with Temporal Reasoner
by: Yang, Xiangpeng, et al.
Published: (2025)
by: Yang, Xiangpeng, et al.
Published: (2025)
UniVideo: Unified Understanding, Generation, and Editing for Videos
by: Wei, Cong, et al.
Published: (2025)
by: Wei, Cong, et al.
Published: (2025)
CamDirector: Towards Long-Term Coherent Video Trajectory Editing
by: Shi, Zhihao, et al.
Published: (2026)
by: Shi, Zhihao, et al.
Published: (2026)
Heterogeneous Graph Transformer for Multiple Tiny Object Tracking in RGB-T Videos
by: Xu, Qingyu, et al.
Published: (2024)
by: Xu, Qingyu, et al.
Published: (2024)
DreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text Guidance
by: Wang, Cong, et al.
Published: (2023)
by: Wang, Cong, et al.
Published: (2023)
Text-based Talking Video Editing with Cascaded Conditional Diffusion
by: Han, Bo, et al.
Published: (2024)
by: Han, Bo, et al.
Published: (2024)
SaMam: Style-aware State Space Model for Arbitrary Image Style Transfer
by: Liu, Hongda, et al.
Published: (2025)
by: Liu, Hongda, et al.
Published: (2025)
Shaping a Stabilized Video by Mitigating Unintended Changes for Concept-Augmented Video Editing
by: Guo, Mingce, et al.
Published: (2024)
by: Guo, Mingce, et al.
Published: (2024)
Action Reimagined: Text-to-Pose Video Editing for Dynamic Human Actions
by: Wang, Lan, et al.
Published: (2024)
by: Wang, Lan, et al.
Published: (2024)
MV-Adapter: Multimodal Video Transfer Learning for Video Text Retrieval
by: Jin, Xiaojie, et al.
Published: (2023)
by: Jin, Xiaojie, et al.
Published: (2023)
Consistent Video Editing as Flow-Driven Image-to-Video Generation
by: Wang, Ge, et al.
Published: (2025)
by: Wang, Ge, et al.
Published: (2025)
COVE: Unleashing the Diffusion Feature Correspondence for Consistent Video Editing
by: Wang, Jiangshan, et al.
Published: (2024)
by: Wang, Jiangshan, et al.
Published: (2024)
CustomVideo: Customizing Text-to-Video Generation with Multiple Subjects
by: Wang, Zhao, et al.
Published: (2024)
by: Wang, Zhao, et al.
Published: (2024)
Text-Video Multi-Grained Integration for Video Moment Montage
by: Yin, Zhihui, et al.
Published: (2024)
by: Yin, Zhihui, et al.
Published: (2024)
TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs
by: Wang, Juntong, et al.
Published: (2025)
by: Wang, Juntong, et al.
Published: (2025)
Streaming Video Diffusion: Online Video Editing with Diffusion Models
by: Chen, Feng, et al.
Published: (2024)
by: Chen, Feng, et al.
Published: (2024)
IC-Effect: Precise and Efficient Video Effects Editing via In-Context Learning
by: Li, Yuanhang, et al.
Published: (2025)
by: Li, Yuanhang, et al.
Published: (2025)
Neural Video Fields Editing
by: Yang, Shuzhou, et al.
Published: (2023)
by: Yang, Shuzhou, et al.
Published: (2023)
Mimir: Improving Video Diffusion Models for Precise Text Understanding
by: Tan, Shuai, et al.
Published: (2024)
by: Tan, Shuai, et al.
Published: (2024)
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis
by: Su, Tongtong, et al.
Published: (2025)
by: Su, Tongtong, et al.
Published: (2025)
DragVideo: Interactive Drag-style Video Editing
by: Deng, Yufan, et al.
Published: (2023)
by: Deng, Yufan, et al.
Published: (2023)
CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
by: Yang, Zhuoyi, et al.
Published: (2024)
by: Yang, Zhuoyi, et al.
Published: (2024)
GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration
by: Huang, Kaiyi, et al.
Published: (2024)
by: Huang, Kaiyi, et al.
Published: (2024)
SketchVideo: Sketch-based Video Generation and Editing
by: Liu, Feng-Lin, et al.
Published: (2025)
by: Liu, Feng-Lin, et al.
Published: (2025)
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models
by: Yi, Hao, et al.
Published: (2024)
by: Yi, Hao, et al.
Published: (2024)
ImVideoEdit: Image-learning Video Editing via 2D Spatial Difference Attention Blocks
by: Xu, Jiayang, et al.
Published: (2026)
by: Xu, Jiayang, et al.
Published: (2026)
DocVideoQA: Towards Comprehensive Understanding of Document-Centric Videos through Question Answering
by: Wang, Haochen, et al.
Published: (2025)
by: Wang, Haochen, et al.
Published: (2025)
PreciseCache: Precise Feature Caching for Efficient and High-fidelity Video Generation
by: Wang, Jiangshan, et al.
Published: (2026)
by: Wang, Jiangshan, et al.
Published: (2026)
LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation
by: Wang, Jiarui, et al.
Published: (2025)
by: Wang, Jiarui, et al.
Published: (2025)
UNIC: Unified In-Context Video Editing
by: Ye, Zixuan, et al.
Published: (2025)
by: Ye, Zixuan, et al.
Published: (2025)
Similar Items
-
FlowDirector: Training-Free Flow Steering for Precise Text-to-Video Editing
by: Li, Guangzhao, et al.
Published: (2025) -
MonoMobility: Zero-Shot 3D Mobility Analysis from Monocular Videos
by: Zhou, Hongyi, et al.
Published: (2025) -
EffiVED:Efficient Video Editing via Text-instruction Diffusion Models
by: Zhang, Zhenghao, et al.
Published: (2024) -
Layout2Scene: 3D Semantic Layout Guided Scene Generation via Geometry and Appearance Diffusion Priors
by: Chen, Minglin, et al.
Published: (2025) -
FastVideoEdit: Leveraging Consistency Models for Efficient Text-to-Video Editing
by: Zhang, Youyuan, et al.
Published: (2024)