OmniTransfer: All-in-one Framework for Spatio-temporal Video Transfer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Pengze, Wu, Yanze, Li, Mengtian, Bai, Xu, Zhao, Songtao, Ye, Fulong, Mou, Chong, Li, Xinghui, Chen, Zhuowei, He, Qian, Gao, Mingyuan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915741537665024
author Zhang, Pengze
Wu, Yanze
Li, Mengtian
Bai, Xu
Zhao, Songtao
Ye, Fulong
Mou, Chong
Li, Xinghui
Chen, Zhuowei
He, Qian
Gao, Mingyuan
author_facet Zhang, Pengze
Wu, Yanze
Li, Mengtian
Bai, Xu
Zhao, Songtao
Ye, Fulong
Mou, Chong
Li, Xinghui
Chen, Zhuowei
He, Qian
Gao, Mingyuan
contents Videos convey richer information than images or text, capturing both spatial and temporal dynamics. However, most existing video customization methods rely on reference images or task-specific temporal priors, failing to fully exploit the rich spatio-temporal information inherent in videos, thereby limiting flexibility and generalization in video generation. To address these limitations, we propose OmniTransfer, a unified framework for spatio-temporal video transfer. It leverages multi-view information across frames to enhance appearance consistency and exploits temporal cues to enable fine-grained temporal control. To unify various video transfer tasks, OmniTransfer incorporates three key designs: Task-aware Positional Bias that adaptively leverages reference video information to improve temporal alignment or appearance consistency; Reference-decoupled Causal Learning separating reference and target branches to enable precise reference transfer while improving efficiency; and Task-adaptive Multimodal Alignment using multimodal semantic guidance to dynamically distinguish and tackle different tasks. Extensive experiments show that OmniTransfer outperforms existing methods in appearance (ID and style) and temporal transfer (camera movement and video effects), while matching pose-guided methods in motion transfer without using pose, establishing a new paradigm for flexible, high-fidelity video generation.
format Preprint
id arxiv_https___arxiv_org_abs_2601_14250
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle OmniTransfer: All-in-one Framework for Spatio-temporal Video Transfer
Zhang, Pengze
Wu, Yanze
Li, Mengtian
Bai, Xu
Zhao, Songtao
Ye, Fulong
Mou, Chong
Li, Xinghui
Chen, Zhuowei
He, Qian
Gao, Mingyuan
Computer Vision and Pattern Recognition
Videos convey richer information than images or text, capturing both spatial and temporal dynamics. However, most existing video customization methods rely on reference images or task-specific temporal priors, failing to fully exploit the rich spatio-temporal information inherent in videos, thereby limiting flexibility and generalization in video generation. To address these limitations, we propose OmniTransfer, a unified framework for spatio-temporal video transfer. It leverages multi-view information across frames to enhance appearance consistency and exploits temporal cues to enable fine-grained temporal control. To unify various video transfer tasks, OmniTransfer incorporates three key designs: Task-aware Positional Bias that adaptively leverages reference video information to improve temporal alignment or appearance consistency; Reference-decoupled Causal Learning separating reference and target branches to enable precise reference transfer while improving efficiency; and Task-adaptive Multimodal Alignment using multimodal semantic guidance to dynamically distinguish and tackle different tasks. Extensive experiments show that OmniTransfer outperforms existing methods in appearance (ID and style) and temporal transfer (camera movement and video effects), while matching pose-guided methods in motion transfer without using pose, establishing a new paradigm for flexible, high-fidelity video generation.
title OmniTransfer: All-in-one Framework for Spatio-temporal Video Transfer
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.14250