Phantom: Subject-consistent video generation via cross-modal alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Lijie, Ma, Tianxiang, Li, Bingchuan, Chen, Zhuowei, Liu, Jiawei, Li, Gen, Zhou, Siyu, He, Qian, Wu, Xinglong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916682338926592
author Liu, Lijie
Ma, Tianxiang
Li, Bingchuan
Chen, Zhuowei
Liu, Jiawei
Li, Gen
Zhou, Siyu
He, Qian
Wu, Xinglong
author_facet Liu, Lijie
Ma, Tianxiang
Li, Bingchuan
Chen, Zhuowei
Liu, Jiawei
Li, Gen
Zhou, Siyu
He, Qian
Wu, Xinglong
contents The continuous development of foundational models for video generation is evolving into various applications, with subject-consistent video generation still in the exploratory stage. We refer to this as Subject-to-Video, which extracts subject elements from reference images and generates subject-consistent videos following textual instructions. We believe that the essence of subject-to-video lies in balancing the dual-modal prompts of text and image, thereby deeply and simultaneously aligning both text and visual content. To this end, we propose Phantom, a unified video generation framework for both single- and multi-subject references. Building on existing text-to-video and image-to-video architectures, we redesign the joint text-image injection model and drive it to learn cross-modal alignment via text-image-video triplet data. The proposed method achieves high-fidelity subject-consistent video generation while addressing issues of image content leakage and multi-subject confusion. Evaluation results indicate that our method outperforms other state-of-the-art closed-source commercial solutions. In particular, we emphasize subject consistency in human generation, covering existing ID-preserving video generation while offering enhanced advantages.
format Preprint
id arxiv_https___arxiv_org_abs_2502_11079
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Phantom: Subject-consistent video generation via cross-modal alignment
Liu, Lijie
Ma, Tianxiang
Li, Bingchuan
Chen, Zhuowei
Liu, Jiawei
Li, Gen
Zhou, Siyu
He, Qian
Wu, Xinglong
Computer Vision and Pattern Recognition
Artificial Intelligence
The continuous development of foundational models for video generation is evolving into various applications, with subject-consistent video generation still in the exploratory stage. We refer to this as Subject-to-Video, which extracts subject elements from reference images and generates subject-consistent videos following textual instructions. We believe that the essence of subject-to-video lies in balancing the dual-modal prompts of text and image, thereby deeply and simultaneously aligning both text and visual content. To this end, we propose Phantom, a unified video generation framework for both single- and multi-subject references. Building on existing text-to-video and image-to-video architectures, we redesign the joint text-image injection model and drive it to learn cross-modal alignment via text-image-video triplet data. The proposed method achieves high-fidelity subject-consistent video generation while addressing issues of image content leakage and multi-subject confusion. Evaluation results indicate that our method outperforms other state-of-the-art closed-source commercial solutions. In particular, we emphasize subject consistency in human generation, covering existing ID-preserving video generation while offering enhanced advantages.
title Phantom: Subject-consistent video generation via cross-modal alignment
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2502.11079