CritiFusion: Semantic Critique and Spectral Alignment for Faithful Text-to-Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, ZhenQi, Ni, TsaiChing, Yang, YuanFu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914226793086976
author Chen, ZhenQi
Ni, TsaiChing
Yang, YuanFu
author_facet Chen, ZhenQi
Ni, TsaiChing
Yang, YuanFu
contents Recent text-to-image diffusion models have achieved remarkable visual fidelity but often struggle with semantic alignment to complex prompts. We introduce CritiFusion, a novel inference-time framework that integrates a multimodal semantic critique mechanism with frequency-domain refinement to improve text-to-image consistency and detail. The proposed CritiCore module leverages a vision-language model and multiple large language models to enrich the prompt context and produce high-level semantic feedback, guiding the diffusion process to better align generated content with the prompt's intent. Additionally, SpecFusion merges intermediate generation states in the spectral domain, injecting coarse structural information while preserving high-frequency details. No additional model training is required. CritiFusion serves as a plug-in refinement stage compatible with existing diffusion backbones. Experiments on standard benchmarks show that our method notably improves human-aligned metrics of text-to-image correspondence and visual quality. CritiFusion consistently boosts performance on human preference scores and aesthetic evaluations, achieving results on par with state-of-the-art reward optimization approaches. Qualitative results further demonstrate superior detail, realism, and prompt fidelity, indicating the effectiveness of our semantic critique and spectral alignment strategy.
format Preprint
id arxiv_https___arxiv_org_abs_2512_22681
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CritiFusion: Semantic Critique and Spectral Alignment for Faithful Text-to-Image Generation
Chen, ZhenQi
Ni, TsaiChing
Yang, YuanFu
Computer Vision and Pattern Recognition
Recent text-to-image diffusion models have achieved remarkable visual fidelity but often struggle with semantic alignment to complex prompts. We introduce CritiFusion, a novel inference-time framework that integrates a multimodal semantic critique mechanism with frequency-domain refinement to improve text-to-image consistency and detail. The proposed CritiCore module leverages a vision-language model and multiple large language models to enrich the prompt context and produce high-level semantic feedback, guiding the diffusion process to better align generated content with the prompt's intent. Additionally, SpecFusion merges intermediate generation states in the spectral domain, injecting coarse structural information while preserving high-frequency details. No additional model training is required. CritiFusion serves as a plug-in refinement stage compatible with existing diffusion backbones. Experiments on standard benchmarks show that our method notably improves human-aligned metrics of text-to-image correspondence and visual quality. CritiFusion consistently boosts performance on human preference scores and aesthetic evaluations, achieving results on par with state-of-the-art reward optimization approaches. Qualitative results further demonstrate superior detail, realism, and prompt fidelity, indicating the effectiveness of our semantic critique and spectral alignment strategy.
title CritiFusion: Semantic Critique and Spectral Alignment for Faithful Text-to-Image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.22681