DLSF: Dual-Layer Synergistic Fusion for High-Fidelity Image Syn-thesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Zhen-Qi, Yang, Yuan-Fu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913948301787136
author Chen, Zhen-Qi
Yang, Yuan-Fu
author_facet Chen, Zhen-Qi
Yang, Yuan-Fu
contents With the rapid advancement of diffusion-based generative models, Stable Diffusion (SD) has emerged as a state-of-the-art framework for high-fidelity im-age synthesis. However, existing SD models suffer from suboptimal feature aggregation, leading to in-complete semantic alignment and loss of fine-grained details, especially in highly textured and complex scenes. To address these limitations, we propose a novel dual-latent integration framework that en-hances feature interactions between the base latent and refined latent representations. Our approach em-ploys a feature concatenation strategy followed by an adaptive fusion module, which can be instantiated as either (i) an Adaptive Global Fusion (AGF) for hier-archical feature harmonization, or (ii) a Dynamic Spatial Fusion (DSF) for spatially-aware refinement. This design enables more effective cross-latent com-munication, preserving both global coherence and local texture fidelity. Our GitHub page: https://anonymous.4open.science/r/MVA2025-22 .
format Preprint
id arxiv_https___arxiv_org_abs_2507_13388
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DLSF: Dual-Layer Synergistic Fusion for High-Fidelity Image Syn-thesis
Chen, Zhen-Qi
Yang, Yuan-Fu
Graphics
With the rapid advancement of diffusion-based generative models, Stable Diffusion (SD) has emerged as a state-of-the-art framework for high-fidelity im-age synthesis. However, existing SD models suffer from suboptimal feature aggregation, leading to in-complete semantic alignment and loss of fine-grained details, especially in highly textured and complex scenes. To address these limitations, we propose a novel dual-latent integration framework that en-hances feature interactions between the base latent and refined latent representations. Our approach em-ploys a feature concatenation strategy followed by an adaptive fusion module, which can be instantiated as either (i) an Adaptive Global Fusion (AGF) for hier-archical feature harmonization, or (ii) a Dynamic Spatial Fusion (DSF) for spatially-aware refinement. This design enables more effective cross-latent com-munication, preserving both global coherence and local texture fidelity. Our GitHub page: https://anonymous.4open.science/r/MVA2025-22 .
title DLSF: Dual-Layer Synergistic Fusion for High-Fidelity Image Syn-thesis
topic Graphics
url https://arxiv.org/abs/2507.13388