MoViE: Mobile Diffusion for Video Editing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866917863006142464 |
|---|---|
| author | Karjauv, Adil Fathima, Noor Lelekas, Ioannis Porikli, Fatih Ghodrati, Amir Habibian, Amirhossein |
| author_facet | Karjauv, Adil Fathima, Noor Lelekas, Ioannis Porikli, Fatih Ghodrati, Amir Habibian, Amirhossein |
| contents | Recent progress in diffusion-based video editing has shown remarkable potential for practical applications. However, these methods remain prohibitively expensive and challenging to deploy on mobile devices. In this study, we introduce a series of optimizations that render mobile video editing feasible. Building upon the existing image editing model, we first optimize its architecture and incorporate a lightweight autoencoder. Subsequently, we extend classifier-free guidance distillation to multiple modalities, resulting in a threefold on-device speedup. Finally, we reduce the number of sampling steps to one by introducing a novel adversarial distillation scheme which preserves the controllability of the editing process. Collectively, these optimizations enable video editing at 12 frames per second on mobile devices, while maintaining high quality. Our results are available at https://qualcomm-ai-research.github.io/mobile-video-editing/ |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_06578 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | MoViE: Mobile Diffusion for Video Editing Karjauv, Adil Fathima, Noor Lelekas, Ioannis Porikli, Fatih Ghodrati, Amir Habibian, Amirhossein Computer Vision and Pattern Recognition Recent progress in diffusion-based video editing has shown remarkable potential for practical applications. However, these methods remain prohibitively expensive and challenging to deploy on mobile devices. In this study, we introduce a series of optimizations that render mobile video editing feasible. Building upon the existing image editing model, we first optimize its architecture and incorporate a lightweight autoencoder. Subsequently, we extend classifier-free guidance distillation to multiple modalities, resulting in a threefold on-device speedup. Finally, we reduce the number of sampling steps to one by introducing a novel adversarial distillation scheme which preserves the controllability of the editing process. Collectively, these optimizations enable video editing at 12 frames per second on mobile devices, while maintaining high quality. Our results are available at https://qualcomm-ai-research.github.io/mobile-video-editing/ |
| title | MoViE: Mobile Diffusion for Video Editing |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2412.06578 |