Zero-Shot Dynamic Concept Personalization with Grid-Based LoRA

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Abdal, Rameen, Patashnik, Or, Deyneka, Ekaterina, Chen, Hao, Siarohin, Aliaksandr, Tulyakov, Sergey, Cohen-Or, Daniel, Aberman, Kfir
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915407652192256
author Abdal, Rameen
Patashnik, Or
Deyneka, Ekaterina
Chen, Hao
Siarohin, Aliaksandr
Tulyakov, Sergey
Cohen-Or, Daniel
Aberman, Kfir
author_facet Abdal, Rameen
Patashnik, Or
Deyneka, Ekaterina
Chen, Hao
Siarohin, Aliaksandr
Tulyakov, Sergey
Cohen-Or, Daniel
Aberman, Kfir
contents Recent advances in text-to-video generation have enabled high-quality synthesis from text and image prompts. While the personalization of dynamic concepts, which capture subject-specific appearance and motion from a single video, is now feasible, most existing methods require per-instance fine-tuning, limiting scalability. We introduce a fully zero-shot framework for dynamic concept personalization in text-to-video models. Our method leverages structured 2x2 video grids that spatially organize input and output pairs, enabling the training of lightweight Grid-LoRA adapters for editing and composition within these grids. At inference, a dedicated Grid Fill module completes partially observed layouts, producing temporally coherent and identity preserving outputs. Once trained, the entire system operates in a single forward pass, generalizing to previously unseen dynamic concepts without any test-time optimization. Extensive experiments demonstrate high-quality and consistent results across a wide range of subjects beyond trained concepts and editing scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2507_17963
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Zero-Shot Dynamic Concept Personalization with Grid-Based LoRA
Abdal, Rameen
Patashnik, Or
Deyneka, Ekaterina
Chen, Hao
Siarohin, Aliaksandr
Tulyakov, Sergey
Cohen-Or, Daniel
Aberman, Kfir
Graphics
Computer Vision and Pattern Recognition
Machine Learning
Recent advances in text-to-video generation have enabled high-quality synthesis from text and image prompts. While the personalization of dynamic concepts, which capture subject-specific appearance and motion from a single video, is now feasible, most existing methods require per-instance fine-tuning, limiting scalability. We introduce a fully zero-shot framework for dynamic concept personalization in text-to-video models. Our method leverages structured 2x2 video grids that spatially organize input and output pairs, enabling the training of lightweight Grid-LoRA adapters for editing and composition within these grids. At inference, a dedicated Grid Fill module completes partially observed layouts, producing temporally coherent and identity preserving outputs. Once trained, the entire system operates in a single forward pass, generalizing to previously unseen dynamic concepts without any test-time optimization. Extensive experiments demonstrate high-quality and consistent results across a wide range of subjects beyond trained concepts and editing scenarios.
title Zero-Shot Dynamic Concept Personalization with Grid-Based LoRA
topic Graphics
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2507.17963