Diffusion Model-Based Video Editing: A Survey
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Wenhao, Tu, Rong-Cheng, Liao, Jingyi, Tao, Dacheng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Diffusion Models, Image Super-Resolution And Everything: A Survey
by: Moser, Brian B., et al.
Published: (2024)
by: Moser, Brian B., et al.
Published: (2024)
Control-A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning
by: Chen, Weifeng, et al.
Published: (2023)
by: Chen, Weifeng, et al.
Published: (2023)
VINCIE: Unlocking In-context Image Editing from Video
by: Qu, Leigang, et al.
Published: (2025)
by: Qu, Leigang, et al.
Published: (2025)
GoodDrag: Towards Good Practices for Drag Editing with Diffusion Models
by: Zhang, Zewei, et al.
Published: (2024)
by: Zhang, Zewei, et al.
Published: (2024)
Meta-CoT: Enhancing Granularity and Generalization in Image Editing
by: Zhang, Shiyi, et al.
Published: (2026)
by: Zhang, Shiyi, et al.
Published: (2026)
IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation
by: Lin, Yuanze, et al.
Published: (2025)
by: Lin, Yuanze, et al.
Published: (2025)
AID: Adapting Image2Video Diffusion Models for Instruction-guided Video Prediction
by: Xing, Zhen, et al.
Published: (2024)
by: Xing, Zhen, et al.
Published: (2024)
PFB-Diff: Progressive Feature Blending Diffusion for Text-driven Image Editing
by: Huang, Wenjing, et al.
Published: (2023)
by: Huang, Wenjing, et al.
Published: (2023)
Neuron Abandoning Attention Flow: Visual Explanation of Dynamics inside CNN Models
by: Liao, Yi, et al.
Published: (2024)
by: Liao, Yi, et al.
Published: (2024)
Towards Multi-Task Multi-Modal Models: A Video Generative Perspective
by: Yu, Lijun
Published: (2024)
by: Yu, Lijun
Published: (2024)
Latent Space Probing for Adult Content Detection in Video Generative Models
by: Khatri, Alizishaan, et al.
Published: (2026)
by: Khatri, Alizishaan, et al.
Published: (2026)
PlanLLM: Video Procedure Planning with Refinable Large Language Models
by: Yang, Dejie, et al.
Published: (2024)
by: Yang, Dejie, et al.
Published: (2024)
Instruction-Guided Editing Controls for Images and Multimedia: A Survey in LLM era
by: Nguyen, Thanh Tam, et al.
Published: (2024)
by: Nguyen, Thanh Tam, et al.
Published: (2024)
RDPM: Solve Diffusion Probabilistic Models via Recurrent Token Prediction
by: Wu, Xiaoping, et al.
Published: (2024)
by: Wu, Xiaoping, et al.
Published: (2024)
State Space Model for New-Generation Network Alternative to Transformers: A Survey
by: Wang, Xiao, et al.
Published: (2024)
by: Wang, Xiao, et al.
Published: (2024)
STIV: Scalable Text and Image Conditioned Video Generation
by: Lin, Zongyu, et al.
Published: (2024)
by: Lin, Zongyu, et al.
Published: (2024)
Multimodal Long Video Modeling Based on Temporal Dynamic Context
by: Hao, Haoran, et al.
Published: (2025)
by: Hao, Haoran, et al.
Published: (2025)
InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions
by: Zhang, Yiyuan, et al.
Published: (2024)
by: Zhang, Yiyuan, et al.
Published: (2024)
TP-Blend: Textual-Prompt Attention Pairing for Precise Object-Style Blending in Diffusion Models
by: Jin, Xin, et al.
Published: (2026)
by: Jin, Xin, et al.
Published: (2026)
David and Goliath: Small One-step Model Beats Large Diffusion with Score Post-training
by: Luo, Weijian, et al.
Published: (2024)
by: Luo, Weijian, et al.
Published: (2024)
MotionCtrl: A Unified and Flexible Motion Controller for Video Generation
by: Wang, Zhouxia, et al.
Published: (2023)
by: Wang, Zhouxia, et al.
Published: (2023)
Bootstrap3D: Improving Multi-view Diffusion Model with Synthetic Data
by: Sun, Zeyi, et al.
Published: (2024)
by: Sun, Zeyi, et al.
Published: (2024)
LayerT2V: A Unified Multi-Layer Video Generation Framework
by: Li, Guangzhao, et al.
Published: (2025)
by: Li, Guangzhao, et al.
Published: (2025)
Zoomed In, Diffused Out: Towards Local Degradation-Aware Multi-Diffusion for Extreme Image Super-Resolution
by: Moser, Brian B., et al.
Published: (2024)
by: Moser, Brian B., et al.
Published: (2024)
Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation
by: Yan, Xin, et al.
Published: (2024)
by: Yan, Xin, et al.
Published: (2024)
VSTAR: Generative Temporal Nursing for Longer Dynamic Video Synthesis
by: Li, Yumeng, et al.
Published: (2024)
by: Li, Yumeng, et al.
Published: (2024)
Reinforcement Learning for Unsupervised Video Summarization with Reward Generator Training
by: Abbasi, Mehryar, et al.
Published: (2024)
by: Abbasi, Mehryar, et al.
Published: (2024)
On-the-fly Modulation for Balanced Multimodal Learning
by: Wei, Yake, et al.
Published: (2024)
by: Wei, Yake, et al.
Published: (2024)
MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance
by: Li, Quanhao, et al.
Published: (2025)
by: Li, Quanhao, et al.
Published: (2025)
FlashMotion: Few-Step Controllable Video Generation with Trajectory Guidance
by: Li, Quanhao, et al.
Published: (2026)
by: Li, Quanhao, et al.
Published: (2026)
Lightning Fast Video Anomaly Detection via Adversarial Knowledge Distillation
by: Croitoru, Florinel-Alin, et al.
Published: (2022)
by: Croitoru, Florinel-Alin, et al.
Published: (2022)
End-to-end Semantic-centric Video-based Multimodal Affective Computing
by: Lin, Ronghao, et al.
Published: (2024)
by: Lin, Ronghao, et al.
Published: (2024)
Adv-KD: Adversarial Knowledge Distillation for Faster Diffusion Sampling
by: Mekonnen, Kidist Amde, et al.
Published: (2024)
by: Mekonnen, Kidist Amde, et al.
Published: (2024)
Video Face Re-Aging: Toward Temporally Consistent Face Re-Aging
by: Muqeet, Abdul, et al.
Published: (2023)
by: Muqeet, Abdul, et al.
Published: (2023)
VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation
by: Zheng, Sixiao, et al.
Published: (2025)
by: Zheng, Sixiao, et al.
Published: (2025)
MAVOS-DD: Multilingual Audio-Video Open-Set Deepfake Detection Benchmark
by: Croitoru, Florinel-Alin, et al.
Published: (2025)
by: Croitoru, Florinel-Alin, et al.
Published: (2025)
Cert-LAS: Toward Certified Model Ownership Verification for Text-to-Image Diffusion Models via Layer-Adaptive Smoothing
by: Qi, Leyi, et al.
Published: (2026)
by: Qi, Leyi, et al.
Published: (2026)
Using AI to Summarize US Presidential Campaign TV Advertisement Videos, 1952-2012
by: Breuer, Adam, et al.
Published: (2025)
by: Breuer, Adam, et al.
Published: (2025)
COMODO: Cross-Modal Video-to-IMU Distillation for Efficient Egocentric Human Activity Recognition
by: Chen, Baiyu, et al.
Published: (2025)
by: Chen, Baiyu, et al.
Published: (2025)
ADS-Edit: A Multimodal Knowledge Editing Dataset for Autonomous Driving Systems
by: Wang, Chenxi, et al.
Published: (2025)
by: Wang, Chenxi, et al.
Published: (2025)
Similar Items
-
Diffusion Models, Image Super-Resolution And Everything: A Survey
by: Moser, Brian B., et al.
Published: (2024) -
Control-A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning
by: Chen, Weifeng, et al.
Published: (2023) -
VINCIE: Unlocking In-context Image Editing from Video
by: Qu, Leigang, et al.
Published: (2025) -
GoodDrag: Towards Good Practices for Drag Editing with Diffusion Models
by: Zhang, Zewei, et al.
Published: (2024) -
Meta-CoT: Enhancing Granularity and Generalization in Image Editing
by: Zhang, Shiyi, et al.
Published: (2026)