Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lv, Zhengyao, Pan, Tianlin, Si, Chenyang, Chen, Zhaoxi, Zuo, Wangmeng, Liu, Ziwei, Wong, Kwan-Yee K.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913954777792512
author Lv, Zhengyao
Pan, Tianlin
Si, Chenyang
Chen, Zhaoxi
Zuo, Wangmeng
Liu, Ziwei
Wong, Kwan-Yee K.
author_facet Lv, Zhengyao
Pan, Tianlin
Si, Chenyang
Chen, Zhaoxi
Zuo, Wangmeng
Liu, Ziwei
Wong, Kwan-Yee K.
contents Multimodal Diffusion Transformers (MM-DiTs) have achieved remarkable progress in text-driven visual generation. However, even state-of-the-art MM-DiT models like FLUX struggle with achieving precise alignment between text prompts and generated content. We identify two key issues in the attention mechanism of MM-DiT, namely 1) the suppression of cross-modal attention due to token imbalance between visual and textual modalities and 2) the lack of timestep-aware attention weighting, which hinder the alignment. To address these issues, we propose \textbf{Temperature-Adjusted Cross-modal Attention (TACA)}, a parameter-efficient method that dynamically rebalances multimodal interactions through temperature scaling and timestep-dependent adjustment. When combined with LoRA fine-tuning, TACA significantly enhances text-image alignment on the T2I-CompBench benchmark with minimal computational overhead. We tested TACA on state-of-the-art models like FLUX and SD3.5, demonstrating its ability to improve image-text alignment in terms of object appearance, attribute binding, and spatial relationships. Our findings highlight the importance of balancing cross-modal attention in improving semantic fidelity in text-to-image diffusion models. Our codes are publicly available at \href{https://github.com/Vchitect/TACA}
format Preprint
id arxiv_https___arxiv_org_abs_2506_07986
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers
Lv, Zhengyao
Pan, Tianlin
Si, Chenyang
Chen, Zhaoxi
Zuo, Wangmeng
Liu, Ziwei
Wong, Kwan-Yee K.
Computer Vision and Pattern Recognition
Multimodal Diffusion Transformers (MM-DiTs) have achieved remarkable progress in text-driven visual generation. However, even state-of-the-art MM-DiT models like FLUX struggle with achieving precise alignment between text prompts and generated content. We identify two key issues in the attention mechanism of MM-DiT, namely 1) the suppression of cross-modal attention due to token imbalance between visual and textual modalities and 2) the lack of timestep-aware attention weighting, which hinder the alignment. To address these issues, we propose \textbf{Temperature-Adjusted Cross-modal Attention (TACA)}, a parameter-efficient method that dynamically rebalances multimodal interactions through temperature scaling and timestep-dependent adjustment. When combined with LoRA fine-tuning, TACA significantly enhances text-image alignment on the T2I-CompBench benchmark with minimal computational overhead. We tested TACA on state-of-the-art models like FLUX and SD3.5, demonstrating its ability to improve image-text alignment in terms of object appearance, attribute binding, and spatial relationships. Our findings highlight the importance of balancing cross-modal attention in improving semantic fidelity in text-to-image diffusion models. Our codes are publicly available at \href{https://github.com/Vchitect/TACA}
title Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.07986