CustomTTT: Motion and Appearance Customized Video Generation via Test-Time Training

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Bi, Xiuli, Lu, Jian, Liu, Bo, Cun, Xiaodong, Zhang, Yong, Li, Weisheng, Xiao, Bin
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916537542115328
author Bi, Xiuli
Lu, Jian
Liu, Bo
Cun, Xiaodong
Zhang, Yong
Li, Weisheng
Xiao, Bin
author_facet Bi, Xiuli
Lu, Jian
Liu, Bo
Cun, Xiaodong
Zhang, Yong
Li, Weisheng
Xiao, Bin
contents Benefiting from large-scale pre-training of text-video pairs, current text-to-video (T2V) diffusion models can generate high-quality videos from the text description. Besides, given some reference images or videos, the parameter-efficient fine-tuning method, i.e. LoRA, can generate high-quality customized concepts, e.g., the specific subject or the motions from a reference video. However, combining the trained multiple concepts from different references into a single network shows obvious artifacts. To this end, we propose CustomTTT, where we can joint custom the appearance and the motion of the given video easily. In detail, we first analyze the prompt influence in the current video diffusion model and find the LoRAs are only needed for the specific layers for appearance and motion customization. Besides, since each LoRA is trained individually, we propose a novel test-time training technique to update parameters after combination utilizing the trained customized models. We conduct detailed experiments to verify the effectiveness of the proposed methods. Our method outperforms several state-of-the-art works in both qualitative and quantitative evaluations.
format Preprint
id arxiv_https___arxiv_org_abs_2412_15646
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CustomTTT: Motion and Appearance Customized Video Generation via Test-Time Training
Bi, Xiuli
Lu, Jian
Liu, Bo
Cun, Xiaodong
Zhang, Yong
Li, Weisheng
Xiao, Bin
Computer Vision and Pattern Recognition
Benefiting from large-scale pre-training of text-video pairs, current text-to-video (T2V) diffusion models can generate high-quality videos from the text description. Besides, given some reference images or videos, the parameter-efficient fine-tuning method, i.e. LoRA, can generate high-quality customized concepts, e.g., the specific subject or the motions from a reference video. However, combining the trained multiple concepts from different references into a single network shows obvious artifacts. To this end, we propose CustomTTT, where we can joint custom the appearance and the motion of the given video easily. In detail, we first analyze the prompt influence in the current video diffusion model and find the LoRAs are only needed for the specific layers for appearance and motion customization. Besides, since each LoRA is trained individually, we propose a novel test-time training technique to update parameters after combination utilizing the trained customized models. We conduct detailed experiments to verify the effectiveness of the proposed methods. Our method outperforms several state-of-the-art works in both qualitative and quantitative evaluations.
title CustomTTT: Motion and Appearance Customized Video Generation via Test-Time Training
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.15646