Saved in:
Bibliographic Details
Main Authors: Mondal, Anindya, Banerjee, Ayan, Nag, Sauradip, Llados, Josep, Zhu, Xiatian, Dutta, Anjan
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2508.16644
Tags: Add Tag
No Tags, Be the first to tag this record!
Table of Contents:
  • Diffusion models excel at photorealistic synthesis but struggle with precise object counts, especially in high-density settings. We introduce COUNTLOOP, a training-free framework that achieves precise instance control through iterative, structured feedback. Our method alternates between synthesis and evaluation: a VLM-based planner generates structured scene layouts, while a VLM-based critic provides explicit feedback on object counts, spatial arrangements, and visual quality to refine the layout iteratively. Instance-driven attention masking and cumulative attention composition further prevent semantic leakage, ensuring clear object separation even in densely occluded scenes. Evaluations on COCO-Count, T2I-CompBench, and two newly introduced high instance benchmarks show that COUNTLOOP reduces counting error by up to 57% and achieves the highest or comparable spatial quality scores across all benchmarks, while maintaining photorealism.