V-Warper: Appearance-Consistent Video Diffusion Personalization via Value Warping

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Hyunkoo, Jang, Wooseok, Yang, Jini, Kim, Taehwan, Kim, Sangoh, Jung, Sangwon, Kim, Seungryong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914200571346944
author Lee, Hyunkoo
Jang, Wooseok
Yang, Jini
Kim, Taehwan
Kim, Sangoh
Jung, Sangwon
Kim, Seungryong
author_facet Lee, Hyunkoo
Jang, Wooseok
Yang, Jini
Kim, Taehwan
Kim, Sangoh
Jung, Sangwon
Kim, Seungryong
contents Video personalization aims to generate videos that faithfully reflect a user-provided subject while following a text prompt. However, existing approaches often rely on heavy video-based finetuning or large-scale video datasets, which impose substantial computational cost and are difficult to scale. Furthermore, they still struggle to maintain fine-grained appearance consistency across frames. To address these limitations, we introduce V-Warper, a training-free coarse-to-fine personalization framework for transformer-based video diffusion models. The framework enhances fine-grained identity fidelity without requiring any additional video training. (1) A lightweight coarse appearance adaptation stage leverages only a small set of reference images, which are already required for the task. This step encodes global subject identity through image-only LoRA and subject-embedding adaptation. (2) A inference-time fine appearance injection stage refines visual fidelity by computing semantic correspondences from RoPE-free mid-layer query--key features. These correspondences guide the warping of appearance-rich value representations into semantically aligned regions of the generation process, with masking ensuring spatial reliability. V-Warper significantly improves appearance fidelity while preserving prompt alignment and motion dynamics, and it achieves these gains efficiently without large-scale video finetuning.
format Preprint
id arxiv_https___arxiv_org_abs_2512_12375
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle V-Warper: Appearance-Consistent Video Diffusion Personalization via Value Warping
Lee, Hyunkoo
Jang, Wooseok
Yang, Jini
Kim, Taehwan
Kim, Sangoh
Jung, Sangwon
Kim, Seungryong
Computer Vision and Pattern Recognition
Video personalization aims to generate videos that faithfully reflect a user-provided subject while following a text prompt. However, existing approaches often rely on heavy video-based finetuning or large-scale video datasets, which impose substantial computational cost and are difficult to scale. Furthermore, they still struggle to maintain fine-grained appearance consistency across frames. To address these limitations, we introduce V-Warper, a training-free coarse-to-fine personalization framework for transformer-based video diffusion models. The framework enhances fine-grained identity fidelity without requiring any additional video training. (1) A lightweight coarse appearance adaptation stage leverages only a small set of reference images, which are already required for the task. This step encodes global subject identity through image-only LoRA and subject-embedding adaptation. (2) A inference-time fine appearance injection stage refines visual fidelity by computing semantic correspondences from RoPE-free mid-layer query--key features. These correspondences guide the warping of appearance-rich value representations into semantically aligned regions of the generation process, with masking ensuring spatial reliability. V-Warper significantly improves appearance fidelity while preserving prompt alignment and motion dynamics, and it achieves these gains efficiently without large-scale video finetuning.
title V-Warper: Appearance-Consistent Video Diffusion Personalization via Value Warping
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.12375