InstanceGen: Image Generation with Instance-level Instructions
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910949975261184 |
|---|---|
| author | Sella, Etai Kleiman, Yanir Averbuch-Elor, Hadar |
| author_facet | Sella, Etai Kleiman, Yanir Averbuch-Elor, Hadar |
| contents | Despite rapid advancements in the capabilities of generative models, pretrained text-to-image models still struggle in capturing the semantics conveyed by complex prompts that compound multiple objects and instance-level attributes. Consequently, we are witnessing growing interests in integrating additional structural constraints, typically in the form of coarse bounding boxes, to better guide the generation process in such challenging cases. In this work, we take the idea of structural guidance a step further by making the observation that contemporary image generation models can directly provide a plausible fine-grained structural initialization. We propose a technique that couples this image-based structural guidance with LLM-based instance-level instructions, yielding output images that adhere to all parts of the text prompt, including object counts, instance-level attributes, and spatial relations between instances. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_05678 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | InstanceGen: Image Generation with Instance-level Instructions Sella, Etai Kleiman, Yanir Averbuch-Elor, Hadar Computer Vision and Pattern Recognition Despite rapid advancements in the capabilities of generative models, pretrained text-to-image models still struggle in capturing the semantics conveyed by complex prompts that compound multiple objects and instance-level attributes. Consequently, we are witnessing growing interests in integrating additional structural constraints, typically in the form of coarse bounding boxes, to better guide the generation process in such challenging cases. In this work, we take the idea of structural guidance a step further by making the observation that contemporary image generation models can directly provide a plausible fine-grained structural initialization. We propose a technique that couples this image-based structural guidance with LLM-based instance-level instructions, yielding output images that adhere to all parts of the text prompt, including object counts, instance-level attributes, and spatial relations between instances. |
| title | InstanceGen: Image Generation with Instance-level Instructions |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2505.05678 |