LoomNet: Enhancing Multi-View Image Generation via Latent Space Weaving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Federico, Giulio, Carrara, Fabio, Gennaro, Claudio, Amato, Giuseppe, Di Benedetto, Marco
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916831762055168
author Federico, Giulio
Carrara, Fabio
Gennaro, Claudio
Amato, Giuseppe
Di Benedetto, Marco
author_facet Federico, Giulio
Carrara, Fabio
Gennaro, Claudio
Amato, Giuseppe
Di Benedetto, Marco
contents Generating consistent multi-view images from a single image remains challenging. Lack of spatial consistency often degrades 3D mesh quality in surface reconstruction. To address this, we propose LoomNet, a novel multi-view diffusion architecture that produces coherent images by applying the same diffusion model multiple times in parallel to collaboratively build and leverage a shared latent space for view consistency. Each viewpoint-specific inference generates an encoding representing its own hypothesis of the novel view from a given camera pose, which is projected onto three orthogonal planes. For each plane, encodings from all views are fused into a single aggregated plane. These aggregated planes are then processed to propagate information and interpolate missing regions, combining the hypotheses into a unified, coherent interpretation. The final latent space is then used to render consistent multi-view images. LoomNet generates 16 high-quality and coherent views in just 15 seconds. In our experiments, LoomNet outperforms state-of-the-art methods on both image quality and reconstruction metrics, also showing creativity by producing diverse, plausible novel views from the same input.
format Preprint
id arxiv_https___arxiv_org_abs_2507_05499
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LoomNet: Enhancing Multi-View Image Generation via Latent Space Weaving
Federico, Giulio
Carrara, Fabio
Gennaro, Claudio
Amato, Giuseppe
Di Benedetto, Marco
Computer Vision and Pattern Recognition
Generating consistent multi-view images from a single image remains challenging. Lack of spatial consistency often degrades 3D mesh quality in surface reconstruction. To address this, we propose LoomNet, a novel multi-view diffusion architecture that produces coherent images by applying the same diffusion model multiple times in parallel to collaboratively build and leverage a shared latent space for view consistency. Each viewpoint-specific inference generates an encoding representing its own hypothesis of the novel view from a given camera pose, which is projected onto three orthogonal planes. For each plane, encodings from all views are fused into a single aggregated plane. These aggregated planes are then processed to propagate information and interpolate missing regions, combining the hypotheses into a unified, coherent interpretation. The final latent space is then used to render consistent multi-view images. LoomNet generates 16 high-quality and coherent views in just 15 seconds. In our experiments, LoomNet outperforms state-of-the-art methods on both image quality and reconstruction metrics, also showing creativity by producing diverse, plausible novel views from the same input.
title LoomNet: Enhancing Multi-View Image Generation via Latent Space Weaving
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.05499