InstanceV: Instance-Level Video Generation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chen, Yuheng, Hu, Teng, Zhang, Jiangning, Xue, Zhucun, Yi, Ran, Ma, Lizhuang
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914173826367488
author Chen, Yuheng
Hu, Teng
Zhang, Jiangning
Xue, Zhucun
Yi, Ran
Ma, Lizhuang
author_facet Chen, Yuheng
Hu, Teng
Zhang, Jiangning
Xue, Zhucun
Yi, Ran
Ma, Lizhuang
contents Recent advances in text-to-video diffusion models have enabled the generation of high-quality videos conditioned on textual descriptions. However, most existing text-to-video models rely solely on textual conditions, lacking general fine-grained controllability over video generation. To address this challenge, we propose InstanceV, a video generation framework that enables i) instance-level control and ii) global semantic consistency. Specifically, with the aid of proposed Instance-aware Masked Cross-Attention mechanism, InstanceV maximizes the utilization of additional instance-level grounding information to generate correctly attributed instances at designated spatial locations. To improve overall consistency, We introduce the Shared Timestep-Adaptive Prompt Enhancement module, which connects local instances with global semantics in a parameter-efficient manner. Furthermore, we incorporate Spatially-Aware Unconditional Guidance during both training and inference to alleviate the disappearance of small instances. Finally, we propose a new benchmark, named InstanceBench, which combines general video quality metrics with instance-aware metrics for more comprehensive evaluation on instance-level video generation. Extensive experiments demonstrate that InstanceV not only achieves remarkable instance-level controllability in video generation, but also outperforms existing state-of-the-art models in both general quality and instance-aware metrics across qualitative and quantitative evaluations.
format Preprint
id arxiv_https___arxiv_org_abs_2511_23146
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle InstanceV: Instance-Level Video Generation
Chen, Yuheng
Hu, Teng
Zhang, Jiangning
Xue, Zhucun
Yi, Ran
Ma, Lizhuang
Computer Vision and Pattern Recognition
Recent advances in text-to-video diffusion models have enabled the generation of high-quality videos conditioned on textual descriptions. However, most existing text-to-video models rely solely on textual conditions, lacking general fine-grained controllability over video generation. To address this challenge, we propose InstanceV, a video generation framework that enables i) instance-level control and ii) global semantic consistency. Specifically, with the aid of proposed Instance-aware Masked Cross-Attention mechanism, InstanceV maximizes the utilization of additional instance-level grounding information to generate correctly attributed instances at designated spatial locations. To improve overall consistency, We introduce the Shared Timestep-Adaptive Prompt Enhancement module, which connects local instances with global semantics in a parameter-efficient manner. Furthermore, we incorporate Spatially-Aware Unconditional Guidance during both training and inference to alleviate the disappearance of small instances. Finally, we propose a new benchmark, named InstanceBench, which combines general video quality metrics with instance-aware metrics for more comprehensive evaluation on instance-level video generation. Extensive experiments demonstrate that InstanceV not only achieves remarkable instance-level controllability in video generation, but also outperforms existing state-of-the-art models in both general quality and instance-aware metrics across qualitative and quantitative evaluations.
title InstanceV: Instance-Level Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.23146