ConsID-Gen: View-Consistent and Identity-Preserving Image-to-Video Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wu, Mingyang, Mishra, Ashirbad, Dey, Soumik, Xing, Shuo, Ravipati, Naveen, Wu, Hansi, Li, Binbin, Tu, Zhengzhong
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915789147209728
author Wu, Mingyang
Mishra, Ashirbad
Dey, Soumik
Xing, Shuo
Ravipati, Naveen
Wu, Hansi
Li, Binbin
Tu, Zhengzhong
author_facet Wu, Mingyang
Mishra, Ashirbad
Dey, Soumik
Xing, Shuo
Ravipati, Naveen
Wu, Hansi
Li, Binbin
Tu, Zhengzhong
contents Image-to-Video generation (I2V) animates a static image into a temporally coherent video sequence following textual instructions, yet preserving fine-grained object identity under changing viewpoints remains a persistent challenge. Unlike text-to-video models, existing I2V pipelines often suffer from appearance drift and geometric distortion, artifacts we attribute to the sparsity of single-view 2D observations and weak cross-modal alignment. Here we address this problem from both data and model perspectives. First, we curate ConsIDVid, a large-scale object-centric dataset built with a scalable pipeline for high-quality, temporally aligned videos, and establish ConsIDVid-Bench, where we present a novel benchmarking and evaluation framework for multi-view consistency using metrics sensitive to subtle geometric and appearance deviations. We further propose ConsID-Gen, a view-assisted I2V generation framework that augments the first frame with unposed auxiliary views and fuses semantic and structural cues via a dual-stream visual-geometric encoder as well as a text-visual connector, yielding unified conditioning for a Diffusion Transformer backbone. Experiments across ConsIDVid-Bench demonstrate that ConsID-Gen consistently outperforms in multiple metrics, with the best overall performance surpassing leading video generation models like Wan2.1 and HunyuanVideo, delivering superior identity fidelity and temporal coherence under challenging real-world scenarios. We will release our model and dataset at https://myangwu.github.io/ConsID-Gen.
format Preprint
id arxiv_https___arxiv_org_abs_2602_10113
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ConsID-Gen: View-Consistent and Identity-Preserving Image-to-Video Generation
Wu, Mingyang
Mishra, Ashirbad
Dey, Soumik
Xing, Shuo
Ravipati, Naveen
Wu, Hansi
Li, Binbin
Tu, Zhengzhong
Computer Vision and Pattern Recognition
Image-to-Video generation (I2V) animates a static image into a temporally coherent video sequence following textual instructions, yet preserving fine-grained object identity under changing viewpoints remains a persistent challenge. Unlike text-to-video models, existing I2V pipelines often suffer from appearance drift and geometric distortion, artifacts we attribute to the sparsity of single-view 2D observations and weak cross-modal alignment. Here we address this problem from both data and model perspectives. First, we curate ConsIDVid, a large-scale object-centric dataset built with a scalable pipeline for high-quality, temporally aligned videos, and establish ConsIDVid-Bench, where we present a novel benchmarking and evaluation framework for multi-view consistency using metrics sensitive to subtle geometric and appearance deviations. We further propose ConsID-Gen, a view-assisted I2V generation framework that augments the first frame with unposed auxiliary views and fuses semantic and structural cues via a dual-stream visual-geometric encoder as well as a text-visual connector, yielding unified conditioning for a Diffusion Transformer backbone. Experiments across ConsIDVid-Bench demonstrate that ConsID-Gen consistently outperforms in multiple metrics, with the best overall performance surpassing leading video generation models like Wan2.1 and HunyuanVideo, delivering superior identity fidelity and temporal coherence under challenging real-world scenarios. We will release our model and dataset at https://myangwu.github.io/ConsID-Gen.
title ConsID-Gen: View-Consistent and Identity-Preserving Image-to-Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.10113