Ar2Can: An Architect and an Artist Leveraging a Canvas for Multi-Human Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Borse, Shubhankar, Pham, Phuc, Farhadzadeh, Farzad, Choi, Seokeon, Nguyen, Phong Ha, Tran, Anh Tuan, Yun, Sungrack, Hayat, Munawar, Porikli, Fatih
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912993005010944
author Borse, Shubhankar
Pham, Phuc
Farhadzadeh, Farzad
Choi, Seokeon
Nguyen, Phong Ha
Tran, Anh Tuan
Yun, Sungrack
Hayat, Munawar
Porikli, Fatih
author_facet Borse, Shubhankar
Pham, Phuc
Farhadzadeh, Farzad
Choi, Seokeon
Nguyen, Phong Ha
Tran, Anh Tuan
Yun, Sungrack
Hayat, Munawar
Porikli, Fatih
contents Despite recent advances in personalized image generation, existing models consistently fail to produce reliable multi-human scenes, often merging or losing facial identity. We present Ar2Can, a novel two-stage framework that disentangles spatial planning from identity rendering for multi-human generation. The Architect predicts structured layouts, specifying where each person should appear. The Artist then synthesizes photorealistic images, guided by a spatially-grounded face matching reward that combines Hungarian spatial alignment with identity similarity. This approach ensures faces are rendered at correct locations and faithfully preserve reference identities. We develop two Architect variants, seamlessly integrated with our diffusion-based Artist model. This is optimized via Group Relative Policy Optimization (GRPO) using compositional rewards for count accuracy, image quality, and identity matching. Evaluated on the MultiHuman-Testbench, Ar2Can achieves substantial improvements in both count accuracy and identity preservation, while maintaining high perceptual quality. Notably, our method achieves these results using primarily synthetic data, without requiring real multi-human images. Project page: https://qualcomm-ai-research.github.io/ar2can/.
format Preprint
id arxiv_https___arxiv_org_abs_2511_22690
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Ar2Can: An Architect and an Artist Leveraging a Canvas for Multi-Human Generation
Borse, Shubhankar
Pham, Phuc
Farhadzadeh, Farzad
Choi, Seokeon
Nguyen, Phong Ha
Tran, Anh Tuan
Yun, Sungrack
Hayat, Munawar
Porikli, Fatih
Computer Vision and Pattern Recognition
Despite recent advances in personalized image generation, existing models consistently fail to produce reliable multi-human scenes, often merging or losing facial identity. We present Ar2Can, a novel two-stage framework that disentangles spatial planning from identity rendering for multi-human generation. The Architect predicts structured layouts, specifying where each person should appear. The Artist then synthesizes photorealistic images, guided by a spatially-grounded face matching reward that combines Hungarian spatial alignment with identity similarity. This approach ensures faces are rendered at correct locations and faithfully preserve reference identities. We develop two Architect variants, seamlessly integrated with our diffusion-based Artist model. This is optimized via Group Relative Policy Optimization (GRPO) using compositional rewards for count accuracy, image quality, and identity matching. Evaluated on the MultiHuman-Testbench, Ar2Can achieves substantial improvements in both count accuracy and identity preservation, while maintaining high perceptual quality. Notably, our method achieves these results using primarily synthetic data, without requiring real multi-human images. Project page: https://qualcomm-ai-research.github.io/ar2can/.
title Ar2Can: An Architect and an Artist Leveraging a Canvas for Multi-Human Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.22690