PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Teng, Yu, Zhentao, Zhou, Zhengguang, Zhang, Jiangning, Zhou, Yuan, Lu, Qinglin, Yi, Ran
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913885874814976
author Hu, Teng
Yu, Zhentao
Zhou, Zhengguang
Zhang, Jiangning
Zhou, Yuan
Lu, Qinglin
Yi, Ran
author_facet Hu, Teng
Yu, Zhentao
Zhou, Zhengguang
Zhang, Jiangning
Zhou, Yuan
Lu, Qinglin
Yi, Ran
contents Despite recent advances in video generation, existing models still lack fine-grained controllability, especially for multi-subject customization with consistent identity and interaction. In this paper, we propose PolyVivid, a multi-subject video customization framework that enables flexible and identity-consistent generation. To establish accurate correspondences between subject images and textual entities, we design a VLLM-based text-image fusion module that embeds visual identities into the textual space for precise grounding. To further enhance identity preservation and subject interaction, we propose a 3D-RoPE-based enhancement module that enables structured bidirectional fusion between text and image embeddings. Moreover, we develop an attention-inherited identity injection module to effectively inject fused identity features into the video generation process, mitigating identity drift. Finally, we construct an MLLM-based data pipeline that combines MLLM-based grounding, segmentation, and a clique-based subject consolidation strategy to produce high-quality multi-subject data, effectively enhancing subject distinction and reducing ambiguity in downstream video generation. Extensive experiments demonstrate that PolyVivid achieves superior performance in identity fidelity, video realism, and subject alignment, outperforming existing open-source and commercial baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2506_07848
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement
Hu, Teng
Yu, Zhentao
Zhou, Zhengguang
Zhang, Jiangning
Zhou, Yuan
Lu, Qinglin
Yi, Ran
Computer Vision and Pattern Recognition
Artificial Intelligence
Despite recent advances in video generation, existing models still lack fine-grained controllability, especially for multi-subject customization with consistent identity and interaction. In this paper, we propose PolyVivid, a multi-subject video customization framework that enables flexible and identity-consistent generation. To establish accurate correspondences between subject images and textual entities, we design a VLLM-based text-image fusion module that embeds visual identities into the textual space for precise grounding. To further enhance identity preservation and subject interaction, we propose a 3D-RoPE-based enhancement module that enables structured bidirectional fusion between text and image embeddings. Moreover, we develop an attention-inherited identity injection module to effectively inject fused identity features into the video generation process, mitigating identity drift. Finally, we construct an MLLM-based data pipeline that combines MLLM-based grounding, segmentation, and a clique-based subject consolidation strategy to produce high-quality multi-subject data, effectively enhancing subject distinction and reducing ambiguity in downstream video generation. Extensive experiments demonstrate that PolyVivid achieves superior performance in identity fidelity, video realism, and subject alignment, outperforming existing open-source and commercial baselines.
title PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2506.07848