Guardado en:
Detalles Bibliográficos
Autores principales: Mondal, Anindya, Banerjee, Ayan, Nag, Sauradip, Llados, Josep, Zhu, Xiatian, Dutta, Anjan
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:https://arxiv.org/abs/2508.16644
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908953012600832
author Mondal, Anindya
Banerjee, Ayan
Nag, Sauradip
Llados, Josep
Zhu, Xiatian
Dutta, Anjan
author_facet Mondal, Anindya
Banerjee, Ayan
Nag, Sauradip
Llados, Josep
Zhu, Xiatian
Dutta, Anjan
contents Diffusion models excel at photorealistic synthesis but struggle with precise object counts, especially in high-density settings. We introduce COUNTLOOP, a training-free framework that achieves precise instance control through iterative, structured feedback. Our method alternates between synthesis and evaluation: a VLM-based planner generates structured scene layouts, while a VLM-based critic provides explicit feedback on object counts, spatial arrangements, and visual quality to refine the layout iteratively. Instance-driven attention masking and cumulative attention composition further prevent semantic leakage, ensuring clear object separation even in densely occluded scenes. Evaluations on COCO-Count, T2I-CompBench, and two newly introduced high instance benchmarks show that COUNTLOOP reduces counting error by up to 57% and achieves the highest or comparable spatial quality scores across all benchmarks, while maintaining photorealism.
format Preprint
id arxiv_https___arxiv_org_abs_2508_16644
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CountLoop: Training-Free High-Instance Image Generation via Iterative Agent Guidance
Mondal, Anindya
Banerjee, Ayan
Nag, Sauradip
Llados, Josep
Zhu, Xiatian
Dutta, Anjan
Computer Vision and Pattern Recognition
Diffusion models excel at photorealistic synthesis but struggle with precise object counts, especially in high-density settings. We introduce COUNTLOOP, a training-free framework that achieves precise instance control through iterative, structured feedback. Our method alternates between synthesis and evaluation: a VLM-based planner generates structured scene layouts, while a VLM-based critic provides explicit feedback on object counts, spatial arrangements, and visual quality to refine the layout iteratively. Instance-driven attention masking and cumulative attention composition further prevent semantic leakage, ensuring clear object separation even in densely occluded scenes. Evaluations on COCO-Count, T2I-CompBench, and two newly introduced high instance benchmarks show that COUNTLOOP reduces counting error by up to 57% and achieves the highest or comparable spatial quality scores across all benchmarks, while maintaining photorealism.
title CountLoop: Training-Free High-Instance Image Generation via Iterative Agent Guidance
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.16644