Scaling Zero-Shot Reference-to-Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Zijian, Liu, Shikun, Liu, Haozhe, Qiu, Haonan, An, Zhaochong, Ren, Weiming, Liu, Zhiheng, Huang, Xiaoke, Ng, Kam Woh, Xie, Tian, Han, Xiao, Cong, Yuren, Li, Hang, Zhu, Chuyan, Patel, Aditya, Xiang, Tao, He, Sen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911306428186624
author Zhou, Zijian
Liu, Shikun
Liu, Haozhe
Qiu, Haonan
An, Zhaochong
Ren, Weiming
Liu, Zhiheng
Huang, Xiaoke
Ng, Kam Woh
Xie, Tian
Han, Xiao
Cong, Yuren
Li, Hang
Zhu, Chuyan
Patel, Aditya
Xiang, Tao
He, Sen
author_facet Zhou, Zijian
Liu, Shikun
Liu, Haozhe
Qiu, Haonan
An, Zhaochong
Ren, Weiming
Liu, Zhiheng
Huang, Xiaoke
Ng, Kam Woh
Xie, Tian
Han, Xiao
Cong, Yuren
Li, Hang
Zhu, Chuyan
Patel, Aditya
Xiang, Tao
He, Sen
contents Reference-to-video (R2V) generation aims to synthesize videos that align with a text prompt while preserving the subject identity from reference images. However, current R2V methods are hindered by the reliance on explicit reference image-video-text triplets, whose construction is highly expensive and difficult to scale. We bypass this bottleneck by introducing Saber, a scalable zero-shot framework that requires no explicit R2V data. Trained exclusively on video-text pairs, Saber employs a masked training strategy and a tailored attention-based model design to learn identity-consistent and reference-aware representations. Mask augmentation techniques are further integrated to mitigate copy-paste artifacts common in reference-to-video generation. Moreover, Saber demonstrates remarkable generalization capabilities across a varying number of references and achieves superior performance on the OpenS2V-Eval benchmark compared to methods trained with R2V data.
format Preprint
id arxiv_https___arxiv_org_abs_2512_06905
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scaling Zero-Shot Reference-to-Video Generation
Zhou, Zijian
Liu, Shikun
Liu, Haozhe
Qiu, Haonan
An, Zhaochong
Ren, Weiming
Liu, Zhiheng
Huang, Xiaoke
Ng, Kam Woh
Xie, Tian
Han, Xiao
Cong, Yuren
Li, Hang
Zhu, Chuyan
Patel, Aditya
Xiang, Tao
He, Sen
Computer Vision and Pattern Recognition
Reference-to-video (R2V) generation aims to synthesize videos that align with a text prompt while preserving the subject identity from reference images. However, current R2V methods are hindered by the reliance on explicit reference image-video-text triplets, whose construction is highly expensive and difficult to scale. We bypass this bottleneck by introducing Saber, a scalable zero-shot framework that requires no explicit R2V data. Trained exclusively on video-text pairs, Saber employs a masked training strategy and a tailored attention-based model design to learn identity-consistent and reference-aware representations. Mask augmentation techniques are further integrated to mitigate copy-paste artifacts common in reference-to-video generation. Moreover, Saber demonstrates remarkable generalization capabilities across a varying number of references and achieves superior performance on the OpenS2V-Eval benchmark compared to methods trained with R2V data.
title Scaling Zero-Shot Reference-to-Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.06905