MoViE: Mobile Diffusion for Video Editing
Fuente:
arXiv
Saved in:
| Main Authors: | Karjauv, Adil, Fathima, Noor, Lelekas, Ioannis, Porikli, Fatih, Ghodrati, Amir, Habibian, Amirhossein |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Neodragon: Mobile Video Generation using Diffusion Transformer
by: Karnewar, Animesh, et al.
Published: (2025)
by: Karnewar, Animesh, et al.
Published: (2025)
Mobile Video Diffusion
by: Yahia, Haitam Ben, et al.
Published: (2024)
by: Yahia, Haitam Ben, et al.
Published: (2024)
Object-Centric Diffusion for Efficient Video Editing
by: Kahatapitiya, Kumara, et al.
Published: (2024)
by: Kahatapitiya, Kumara, et al.
Published: (2024)
Clockwork Diffusion: Efficient Generation With Model-Step Distillation
by: Habibian, Amirhossein, et al.
Published: (2023)
by: Habibian, Amirhossein, et al.
Published: (2023)
PyramidalWan: On Making Pretrained Video Model Pyramidal for Efficient Inference
by: Korzhenkov, Denis, et al.
Published: (2026)
by: Korzhenkov, Denis, et al.
Published: (2026)
ReHyAt: Recurrent Hybrid Attention for Video Diffusion Transformers
by: Ghafoorian, Mohsen, et al.
Published: (2026)
by: Ghafoorian, Mohsen, et al.
Published: (2026)
Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformer
by: Ghafoorian, Mohsen, et al.
Published: (2025)
by: Ghafoorian, Mohsen, et al.
Published: (2025)
MoAlign: Motion-Centric Representation Alignment for Video Diffusion Models
by: Bhowmik, Aritra, et al.
Published: (2025)
by: Bhowmik, Aritra, et al.
Published: (2025)
Enhancing Novel View Synthesis via Geometry Grounded Set Diffusion
by: Zanjani, Farhad G., et al.
Published: (2026)
by: Zanjani, Farhad G., et al.
Published: (2026)
MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing
by: Kadambi, Shreya, et al.
Published: (2025)
by: Kadambi, Shreya, et al.
Published: (2025)
Multi-Scale Local Speculative Decoding for Image Generation
by: Peruzzo, Elia, et al.
Published: (2026)
by: Peruzzo, Elia, et al.
Published: (2026)
Segmentation-Free Guidance for Text-to-Image Diffusion Models
by: Azarian, Kambiz, et al.
Published: (2024)
by: Azarian, Kambiz, et al.
Published: (2024)
Controllable 3D Placement of Objects with Scene-Aware Diffusion Models
by: Omran, Mohamed, et al.
Published: (2025)
by: Omran, Mohamed, et al.
Published: (2025)
DySS: Dynamic Queries and State-Space Learning for Efficient 3D Object Detection from Multi-Camera Videos
by: Yasarla, Rajeev, et al.
Published: (2025)
by: Yasarla, Rajeev, et al.
Published: (2025)
LaFAM: Unsupervised Feature Attribution with Label-free Activation Maps
by: Karjauv, Aray, et al.
Published: (2024)
by: Karjauv, Aray, et al.
Published: (2024)
Scene-Aware Location Modeling for Data Augmentation in Automotive Object Detection
by: Petersen, Jens, et al.
Published: (2025)
by: Petersen, Jens, et al.
Published: (2025)
MAMo: Leveraging Memory and Attention for Monocular Video Depth Estimation
by: Yasarla, Rajeev, et al.
Published: (2023)
by: Yasarla, Rajeev, et al.
Published: (2023)
CoReDiT: Spatial Coherence-Guided Token Pruning and Reconstruction for Efficient Diffusion Transformers
by: Li, Zhuojin, et al.
Published: (2026)
by: Li, Zhuojin, et al.
Published: (2026)
Resolving the Identity Crisis in Text-to-Image Generation
by: Borse, Shubhankar, et al.
Published: (2025)
by: Borse, Shubhankar, et al.
Published: (2025)
H3O: Hyper-Efficient 3D Occupancy Prediction with Heterogeneous Supervision
by: Shi, Yunxiao, et al.
Published: (2025)
by: Shi, Yunxiao, et al.
Published: (2025)
LoRA-X: Bridging Foundation Models with Training-Free Cross-Model Adaptation
by: Farhadzadeh, Farzad, et al.
Published: (2025)
by: Farhadzadeh, Farzad, et al.
Published: (2025)
Hybrid Gaussian Splatting for Novel Urban View Synthesis
by: Omran, Mohamed, et al.
Published: (2025)
by: Omran, Mohamed, et al.
Published: (2025)
MobileViCLIP: An Efficient Video-Text Model for Mobile Devices
by: Yang, Min, et al.
Published: (2025)
by: Yang, Min, et al.
Published: (2025)
DDIL: Diversity Enhancing Diffusion Distillation With Imitation Learning
by: Garrepalli, Risheek, et al.
Published: (2024)
by: Garrepalli, Risheek, et al.
Published: (2024)
Gated Relational Alignment via Confidence-based Distillation for Efficient VLMs
by: Chen, Yanlong, et al.
Published: (2026)
by: Chen, Yanlong, et al.
Published: (2026)
ViViD: Video Virtual Try-on using Diffusion Models
by: Fang, Zixun, et al.
Published: (2024)
by: Fang, Zixun, et al.
Published: (2024)
DeCoTR: Enhancing Depth Completion with 2D and 3D Attentions
by: Shi, Yunxiao, et al.
Published: (2024)
by: Shi, Yunxiao, et al.
Published: (2024)
Improving Optical Flow and Stereo Depth Estimation by Leveraging Uncertainty-Based Learning Difficulties
by: Jeong, Jisoo, et al.
Published: (2025)
by: Jeong, Jisoo, et al.
Published: (2025)
EdgeRelight360: Text-Conditioned 360-Degree HDR Image Generation for Real-Time On-Device Video Portrait Relighting
by: Lin, Min-Hui, et al.
Published: (2024)
by: Lin, Min-Hui, et al.
Published: (2024)
InsViE-1M: Effective Instruction-based Video Editing with Elaborate Dataset Construction
by: Wu, Yuhui, et al.
Published: (2025)
by: Wu, Yuhui, et al.
Published: (2025)
Imagining the Unseen: Generative Location Modeling for Object Placement
by: Yun, Jooyeol, et al.
Published: (2024)
by: Yun, Jooyeol, et al.
Published: (2024)
HexaGen3D: StableDiffusion is just one step away from Fast and Diverse Text-to-3D Generation
by: Mercier, Antoine, et al.
Published: (2024)
by: Mercier, Antoine, et al.
Published: (2024)
FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video Understanding
by: Cho, Janghoon, et al.
Published: (2025)
by: Cho, Janghoon, et al.
Published: (2025)
Streaming Video Diffusion: Online Video Editing with Diffusion Models
by: Chen, Feng, et al.
Published: (2024)
by: Chen, Feng, et al.
Published: (2024)
Sort-free Gaussian Splatting via Weighted Sum Rendering
by: Hou, Qiqi, et al.
Published: (2024)
by: Hou, Qiqi, et al.
Published: (2024)
Temporally Consistent Object Editing in Videos using Extended Attention
by: Zamani, AmirHossein, et al.
Published: (2024)
by: Zamani, AmirHossein, et al.
Published: (2024)
RepViT: Revisiting Mobile CNN From ViT Perspective
by: Wang, Ao, et al.
Published: (2023)
by: Wang, Ao, et al.
Published: (2023)
ToSA: Token Selective Attention for Efficient Vision Transformers
by: Singh, Manish Kumar, et al.
Published: (2024)
by: Singh, Manish Kumar, et al.
Published: (2024)
Planar Gaussian Splatting
by: Zanjani, Farhad G., et al.
Published: (2024)
by: Zanjani, Farhad G., et al.
Published: (2024)
Neural Mesh Fusion: Unsupervised 3D Planar Surface Understanding
by: Zanjani, Farhad G., et al.
Published: (2024)
by: Zanjani, Farhad G., et al.
Published: (2024)
Similar Items
-
Neodragon: Mobile Video Generation using Diffusion Transformer
by: Karnewar, Animesh, et al.
Published: (2025) -
Mobile Video Diffusion
by: Yahia, Haitam Ben, et al.
Published: (2024) -
Object-Centric Diffusion for Efficient Video Editing
by: Kahatapitiya, Kumara, et al.
Published: (2024) -
Clockwork Diffusion: Efficient Generation With Model-Step Distillation
by: Habibian, Amirhossein, et al.
Published: (2023) -
PyramidalWan: On Making Pretrained Video Model Pyramidal for Efficient Inference
by: Korzhenkov, Denis, et al.
Published: (2026)