Guardado en:
Detalles Bibliográficos
Autores principales: Peng, Wenxuan, Hariharan, Bharath, Averbuch-Elor, Hadar
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:https://arxiv.org/abs/2605.23178
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917522991742976
author Peng, Wenxuan
Hariharan, Bharath
Averbuch-Elor, Hadar
author_facet Peng, Wenxuan
Hariharan, Bharath
Averbuch-Elor, Hadar
contents Despite recent progress, text-to-image models still struggle to generate semantically diverse and compositionally accurate multi-person interaction scenes, often collapsing to repetitive layouts, stereotypical poses, and poorly grounded interactions. In this work, we bridge this gap by introducing a dual pose-image representation that brings person-centric structural priors into pretrained diffusion transformers. Our model jointly predicts a 2D pose visualization image and its corresponding RGB image, enabling structure and appearance to co-evolve during learning. At its core, a cross-modal alignment scheme binds text, pose, and image representations, ensuring consistent grounding across modalities. Furthermore, we design an iterative scene construction scheme, progressively generating complex multi-human interactions while effectively decomposing the overall generation complexity. Extensive experiments demonstrate that our method substantially improves prompt alignment and scene diversity in multi-person image generation.
format Preprint
id arxiv_https___arxiv_org_abs_2605_23178
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Composing People Together: Iterative Pose-Image Generation for Multi-Person Interaction Scenes
Peng, Wenxuan
Hariharan, Bharath
Averbuch-Elor, Hadar
Computer Vision and Pattern Recognition
Despite recent progress, text-to-image models still struggle to generate semantically diverse and compositionally accurate multi-person interaction scenes, often collapsing to repetitive layouts, stereotypical poses, and poorly grounded interactions. In this work, we bridge this gap by introducing a dual pose-image representation that brings person-centric structural priors into pretrained diffusion transformers. Our model jointly predicts a 2D pose visualization image and its corresponding RGB image, enabling structure and appearance to co-evolve during learning. At its core, a cross-modal alignment scheme binds text, pose, and image representations, ensuring consistent grounding across modalities. Furthermore, we design an iterative scene construction scheme, progressively generating complex multi-human interactions while effectively decomposing the overall generation complexity. Extensive experiments demonstrate that our method substantially improves prompt alignment and scene diversity in multi-person image generation.
title Composing People Together: Iterative Pose-Image Generation for Multi-Person Interaction Scenes
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.23178