First Frame Is the Place to Go for Video Content Customization

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chen, Jingxi, Li, Zongxia, Liu, Zhichao, Shi, Guangyao, Wu, Xiyang, Liu, Fuxiao, Fermuller, Cornelia, Feng, Brandon Y., Aloimonos, Yiannis
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917357474021376
author Chen, Jingxi
Li, Zongxia
Liu, Zhichao
Shi, Guangyao
Wu, Xiyang
Liu, Fuxiao
Fermuller, Cornelia
Feng, Brandon Y.
Aloimonos, Yiannis
author_facet Chen, Jingxi
Li, Zongxia
Liu, Zhichao
Shi, Guangyao
Wu, Xiyang
Liu, Fuxiao
Fermuller, Cornelia
Feng, Brandon Y.
Aloimonos, Yiannis
contents What role does the first frame play in video generation models? Traditionally, it's viewed as the spatial-temporal starting point of a video, merely a seed for subsequent animation. In this work, we reveal a fundamentally different perspective: video models implicitly treat the first frame as a conceptual memory buffer that stores visual entities for later reuse during generation. Leveraging this insight, we show that it's possible to achieve robust and generalized video content customization in diverse scenarios, using only 20-50 training examples without architectural changes or large-scale finetuning. This unveils a powerful, overlooked capability of video generation models for reference-based video customization.
format Preprint
id arxiv_https___arxiv_org_abs_2511_15700
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle First Frame Is the Place to Go for Video Content Customization
Chen, Jingxi
Li, Zongxia
Liu, Zhichao
Shi, Guangyao
Wu, Xiyang
Liu, Fuxiao
Fermuller, Cornelia
Feng, Brandon Y.
Aloimonos, Yiannis
Computer Vision and Pattern Recognition
What role does the first frame play in video generation models? Traditionally, it's viewed as the spatial-temporal starting point of a video, merely a seed for subsequent animation. In this work, we reveal a fundamentally different perspective: video models implicitly treat the first frame as a conceptual memory buffer that stores visual entities for later reuse during generation. Leveraging this insight, we show that it's possible to achieve robust and generalized video content customization in diverse scenarios, using only 20-50 training examples without architectural changes or large-scale finetuning. This unveils a powerful, overlooked capability of video generation models for reference-based video customization.
title First Frame Is the Place to Go for Video Content Customization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.15700