Scaling Zero-Shot Reference-to-Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911306428186624 |
|---|---|
| author | Zhou, Zijian Liu, Shikun Liu, Haozhe Qiu, Haonan An, Zhaochong Ren, Weiming Liu, Zhiheng Huang, Xiaoke Ng, Kam Woh Xie, Tian Han, Xiao Cong, Yuren Li, Hang Zhu, Chuyan Patel, Aditya Xiang, Tao He, Sen |
| author_facet | Zhou, Zijian Liu, Shikun Liu, Haozhe Qiu, Haonan An, Zhaochong Ren, Weiming Liu, Zhiheng Huang, Xiaoke Ng, Kam Woh Xie, Tian Han, Xiao Cong, Yuren Li, Hang Zhu, Chuyan Patel, Aditya Xiang, Tao He, Sen |
| contents | Reference-to-video (R2V) generation aims to synthesize videos that align with a text prompt while preserving the subject identity from reference images. However, current R2V methods are hindered by the reliance on explicit reference image-video-text triplets, whose construction is highly expensive and difficult to scale. We bypass this bottleneck by introducing Saber, a scalable zero-shot framework that requires no explicit R2V data. Trained exclusively on video-text pairs, Saber employs a masked training strategy and a tailored attention-based model design to learn identity-consistent and reference-aware representations. Mask augmentation techniques are further integrated to mitigate copy-paste artifacts common in reference-to-video generation. Moreover, Saber demonstrates remarkable generalization capabilities across a varying number of references and achieves superior performance on the OpenS2V-Eval benchmark compared to methods trained with R2V data. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_06905 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Scaling Zero-Shot Reference-to-Video Generation Zhou, Zijian Liu, Shikun Liu, Haozhe Qiu, Haonan An, Zhaochong Ren, Weiming Liu, Zhiheng Huang, Xiaoke Ng, Kam Woh Xie, Tian Han, Xiao Cong, Yuren Li, Hang Zhu, Chuyan Patel, Aditya Xiang, Tao He, Sen Computer Vision and Pattern Recognition Reference-to-video (R2V) generation aims to synthesize videos that align with a text prompt while preserving the subject identity from reference images. However, current R2V methods are hindered by the reliance on explicit reference image-video-text triplets, whose construction is highly expensive and difficult to scale. We bypass this bottleneck by introducing Saber, a scalable zero-shot framework that requires no explicit R2V data. Trained exclusively on video-text pairs, Saber employs a masked training strategy and a tailored attention-based model design to learn identity-consistent and reference-aware representations. Mask augmentation techniques are further integrated to mitigate copy-paste artifacts common in reference-to-video generation. Moreover, Saber demonstrates remarkable generalization capabilities across a varying number of references and achieves superior performance on the OpenS2V-Eval benchmark compared to methods trained with R2V data. |
| title | Scaling Zero-Shot Reference-to-Video Generation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2512.06905 |