RefAlign: Representation Alignment for Reference-to-Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Lei, Song, YuXin, Wu, Ge, Feng, Haocheng, Zhou, Hang, Wang, Jingdong, Wang, Yaxing, Yang, jian
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910076128722944
author Wang, Lei
Song, YuXin
Wu, Ge
Feng, Haocheng
Zhou, Hang
Wang, Jingdong
Wang, Yaxing
Yang, jian
author_facet Wang, Lei
Song, YuXin
Wu, Ge
Feng, Haocheng
Zhou, Hang
Wang, Jingdong
Wang, Yaxing
Yang, jian
contents Reference-to-video (R2V) generation is a controllable video synthesis paradigm that constrains the generation process using both text prompts and reference images, enabling applications such as personalized advertising and virtual try-on. In practice, existing R2V methods typically introduce additional high-level semantic or cross-modal features alongside the VAE latent representation of the reference image and jointly feed them into the diffusion Transformer (DiT). These auxiliary representations provide semantic guidance and act as implicit alignment signals, which can partially alleviate pixel-level information leakage in the VAE latent space. However, they may still struggle to address copy--paste artifacts and multi-subject confusion caused by modality mismatch across heterogeneous encoder features. In this paper, we propose RefAlign, a representation alignment framework that explicitly aligns DiT reference-branch features to the semantic space of a visual foundation model (VFM). The core of RefAlign is a reference alignment loss that pulls the reference features and VFM features of the same subject closer to improve identity consistency, while pushing apart the corresponding features of different subjects to enhance semantic discriminability. This simple yet effective strategy is applied only during training, incurring no inference-time overhead, and achieves a better balance between text controllability and reference fidelity. Extensive experiments on the OpenS2V-Eval benchmark demonstrate that RefAlign outperforms current state-of-the-art methods in TotalScore, validating the effectiveness of explicit reference alignment for R2V tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2603_25743
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RefAlign: Representation Alignment for Reference-to-Video Generation
Wang, Lei
Song, YuXin
Wu, Ge
Feng, Haocheng
Zhou, Hang
Wang, Jingdong
Wang, Yaxing
Yang, jian
Computer Vision and Pattern Recognition
Reference-to-video (R2V) generation is a controllable video synthesis paradigm that constrains the generation process using both text prompts and reference images, enabling applications such as personalized advertising and virtual try-on. In practice, existing R2V methods typically introduce additional high-level semantic or cross-modal features alongside the VAE latent representation of the reference image and jointly feed them into the diffusion Transformer (DiT). These auxiliary representations provide semantic guidance and act as implicit alignment signals, which can partially alleviate pixel-level information leakage in the VAE latent space. However, they may still struggle to address copy--paste artifacts and multi-subject confusion caused by modality mismatch across heterogeneous encoder features. In this paper, we propose RefAlign, a representation alignment framework that explicitly aligns DiT reference-branch features to the semantic space of a visual foundation model (VFM). The core of RefAlign is a reference alignment loss that pulls the reference features and VFM features of the same subject closer to improve identity consistency, while pushing apart the corresponding features of different subjects to enhance semantic discriminability. This simple yet effective strategy is applied only during training, incurring no inference-time overhead, and achieves a better balance between text controllability and reference fidelity. Extensive experiments on the OpenS2V-Eval benchmark demonstrate that RefAlign outperforms current state-of-the-art methods in TotalScore, validating the effectiveness of explicit reference alignment for R2V tasks.
title RefAlign: Representation Alignment for Reference-to-Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.25743