Zero-Shot Dynamic Concept Personalization with Grid-Based LoRA
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915407652192256 |
|---|---|
| author | Abdal, Rameen Patashnik, Or Deyneka, Ekaterina Chen, Hao Siarohin, Aliaksandr Tulyakov, Sergey Cohen-Or, Daniel Aberman, Kfir |
| author_facet | Abdal, Rameen Patashnik, Or Deyneka, Ekaterina Chen, Hao Siarohin, Aliaksandr Tulyakov, Sergey Cohen-Or, Daniel Aberman, Kfir |
| contents | Recent advances in text-to-video generation have enabled high-quality synthesis from text and image prompts. While the personalization of dynamic concepts, which capture subject-specific appearance and motion from a single video, is now feasible, most existing methods require per-instance fine-tuning, limiting scalability. We introduce a fully zero-shot framework for dynamic concept personalization in text-to-video models. Our method leverages structured 2x2 video grids that spatially organize input and output pairs, enabling the training of lightweight Grid-LoRA adapters for editing and composition within these grids. At inference, a dedicated Grid Fill module completes partially observed layouts, producing temporally coherent and identity preserving outputs. Once trained, the entire system operates in a single forward pass, generalizing to previously unseen dynamic concepts without any test-time optimization. Extensive experiments demonstrate high-quality and consistent results across a wide range of subjects beyond trained concepts and editing scenarios. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_17963 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Zero-Shot Dynamic Concept Personalization with Grid-Based LoRA Abdal, Rameen Patashnik, Or Deyneka, Ekaterina Chen, Hao Siarohin, Aliaksandr Tulyakov, Sergey Cohen-Or, Daniel Aberman, Kfir Graphics Computer Vision and Pattern Recognition Machine Learning Recent advances in text-to-video generation have enabled high-quality synthesis from text and image prompts. While the personalization of dynamic concepts, which capture subject-specific appearance and motion from a single video, is now feasible, most existing methods require per-instance fine-tuning, limiting scalability. We introduce a fully zero-shot framework for dynamic concept personalization in text-to-video models. Our method leverages structured 2x2 video grids that spatially organize input and output pairs, enabling the training of lightweight Grid-LoRA adapters for editing and composition within these grids. At inference, a dedicated Grid Fill module completes partially observed layouts, producing temporally coherent and identity preserving outputs. Once trained, the entire system operates in a single forward pass, generalizing to previously unseen dynamic concepts without any test-time optimization. Extensive experiments demonstrate high-quality and consistent results across a wide range of subjects beyond trained concepts and editing scenarios. |
| title | Zero-Shot Dynamic Concept Personalization with Grid-Based LoRA |
| topic | Graphics Computer Vision and Pattern Recognition Machine Learning |
| url | https://arxiv.org/abs/2507.17963 |