AdaptaGen: Domain-Specific Image Generation through Hierarchical Semantic Optimization Framework

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhang, Suoxiang, Li, Xiaxi, Chang, Hongrui, Hou, Zhuoyan, Wu, Guoxin, Ji, Ronghua
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916832035733504
author Zhang, Suoxiang
Li, Xiaxi
Chang, Hongrui
Hou, Zhuoyan
Wu, Guoxin
Ji, Ronghua
author_facet Zhang, Suoxiang
Li, Xiaxi
Chang, Hongrui
Hou, Zhuoyan
Wu, Guoxin
Ji, Ronghua
contents Domain-specific image generation aims to produce high-quality visual content for specialized fields while ensuring semantic accuracy and detail fidelity. However, existing methods exhibit two critical limitations: First, current approaches address prompt engineering and model adaptation separately, overlooking the inherent dependence between semantic understanding and visual representation in specialized domains. Second, these techniques inadequately incorporate domain-specific semantic constraints during content synthesis, resulting in generation outcomes that exhibit hallucinations and semantic deviations. To tackle these issues, we propose AdaptaGen, a hierarchical semantic optimization framework that integrates matrix-based prompt optimization with multi-perspective understanding, capturing comprehensive semantic relationships from both global and local perspectives. To mitigate hallucinations in specialized domains, we design a cross-modal adaptation mechanism, which, when combined with intelligent content synthesis, enables preserving core thematic elements while incorporating diverse details across images. Additionally, we introduce a two-phase caption semantic transformation during the generation phase. This approach maintains semantic coherence while enhancing visual diversity, ensuring the generated images adhere to domain-specific constraints. Experimental results confirm our approach's effectiveness, with our framework achieving superior performance across 40 categories from diverse datasets using only 16 images per category, demonstrating significant improvements in image quality, diversity, and semantic consistency.
format Preprint
id arxiv_https___arxiv_org_abs_2507_05621
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AdaptaGen: Domain-Specific Image Generation through Hierarchical Semantic Optimization Framework
Zhang, Suoxiang
Li, Xiaxi
Chang, Hongrui
Hou, Zhuoyan
Wu, Guoxin
Ji, Ronghua
Computer Vision and Pattern Recognition
Multimedia
Domain-specific image generation aims to produce high-quality visual content for specialized fields while ensuring semantic accuracy and detail fidelity. However, existing methods exhibit two critical limitations: First, current approaches address prompt engineering and model adaptation separately, overlooking the inherent dependence between semantic understanding and visual representation in specialized domains. Second, these techniques inadequately incorporate domain-specific semantic constraints during content synthesis, resulting in generation outcomes that exhibit hallucinations and semantic deviations. To tackle these issues, we propose AdaptaGen, a hierarchical semantic optimization framework that integrates matrix-based prompt optimization with multi-perspective understanding, capturing comprehensive semantic relationships from both global and local perspectives. To mitigate hallucinations in specialized domains, we design a cross-modal adaptation mechanism, which, when combined with intelligent content synthesis, enables preserving core thematic elements while incorporating diverse details across images. Additionally, we introduce a two-phase caption semantic transformation during the generation phase. This approach maintains semantic coherence while enhancing visual diversity, ensuring the generated images adhere to domain-specific constraints. Experimental results confirm our approach's effectiveness, with our framework achieving superior performance across 40 categories from diverse datasets using only 16 images per category, demonstrating significant improvements in image quality, diversity, and semantic consistency.
title AdaptaGen: Domain-Specific Image Generation through Hierarchical Semantic Optimization Framework
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2507.05621