CEIDM: A Controlled Entity and Interaction Diffusion Model for Enhanced Text-to-Image Generation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yang, Mingyue, Shi, Dianxi, Zhou, Jialu, Wei, Xinyu, Li, Leqian, Yang, Shaowu, Qiu, Chunping
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911119864496128
author Yang, Mingyue
Shi, Dianxi
Zhou, Jialu
Wei, Xinyu
Li, Leqian
Yang, Shaowu
Qiu, Chunping
author_facet Yang, Mingyue
Shi, Dianxi
Zhou, Jialu
Wei, Xinyu
Li, Leqian
Yang, Shaowu
Qiu, Chunping
contents In Text-to-Image (T2I) generation, the complexity of entities and their intricate interactions pose a significant challenge for T2I method based on diffusion model: how to effectively control entity and their interactions to produce high-quality images. To address this, we propose CEIDM, a image generation method based on diffusion model with dual controls for entity and interaction. First, we propose an entity interactive relationships mining approach based on Large Language Models (LLMs), extracting reasonable and rich implicit interactive relationships through chain of thought to guide diffusion models to generate high-quality images that are closer to realistic logic and have more reasonable interactive relationships. Furthermore, We propose an interactive action clustering and offset method to cluster and offset the interactive action features contained in each text prompts. By constructing global and local bidirectional offsets, we enhance semantic understanding and detail supplementation of original actions, making the model's understanding of the concept of interactive "actions" more accurate and generating images with more accurate interactive actions. Finally, we design an entity control network which generates masks with entity semantic guidance, then leveraging multi-scale convolutional network to enhance entity feature and dynamic network to fuse feature. It effectively controls entities and significantly improves image quality. Experiments show that the proposed CEIDM method is better than the most representative existing methods in both entity control and their interaction control.
format Preprint
id arxiv_https___arxiv_org_abs_2508_17760
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CEIDM: A Controlled Entity and Interaction Diffusion Model for Enhanced Text-to-Image Generation
Yang, Mingyue
Shi, Dianxi
Zhou, Jialu
Wei, Xinyu
Li, Leqian
Yang, Shaowu
Qiu, Chunping
Computer Vision and Pattern Recognition
Computation and Language
In Text-to-Image (T2I) generation, the complexity of entities and their intricate interactions pose a significant challenge for T2I method based on diffusion model: how to effectively control entity and their interactions to produce high-quality images. To address this, we propose CEIDM, a image generation method based on diffusion model with dual controls for entity and interaction. First, we propose an entity interactive relationships mining approach based on Large Language Models (LLMs), extracting reasonable and rich implicit interactive relationships through chain of thought to guide diffusion models to generate high-quality images that are closer to realistic logic and have more reasonable interactive relationships. Furthermore, We propose an interactive action clustering and offset method to cluster and offset the interactive action features contained in each text prompts. By constructing global and local bidirectional offsets, we enhance semantic understanding and detail supplementation of original actions, making the model's understanding of the concept of interactive "actions" more accurate and generating images with more accurate interactive actions. Finally, we design an entity control network which generates masks with entity semantic guidance, then leveraging multi-scale convolutional network to enhance entity feature and dynamic network to fuse feature. It effectively controls entities and significantly improves image quality. Experiments show that the proposed CEIDM method is better than the most representative existing methods in both entity control and their interaction control.
title CEIDM: A Controlled Entity and Interaction Diffusion Model for Enhanced Text-to-Image Generation
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2508.17760