RFMI: Estimating Mutual Information on Rectified Flow for Text-to-Image Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Chao, Franzese, Giulio, Finamore, Alessandro, Michiardi, Pietro
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908273653841920
author Wang, Chao
Franzese, Giulio
Finamore, Alessandro
Michiardi, Pietro
author_facet Wang, Chao
Franzese, Giulio
Finamore, Alessandro
Michiardi, Pietro
contents Rectified Flow (RF) models trained with a Flow matching framework have achieved state-of-the-art performance on Text-to-Image (T2I) conditional generation. Yet, multiple benchmarks show that synthetic images can still suffer from poor alignment with the prompt, i.e., images show wrong attribute binding, subject positioning, numeracy, etc. While the literature offers many methods to improve T2I alignment, they all consider only Diffusion Models, and require auxiliary datasets, scoring models, and linguistic analysis of the prompt. In this paper we aim to address these gaps. First, we introduce RFMI, a novel Mutual Information (MI) estimator for RF models that uses the pre-trained model itself for the MI estimation. Then, we investigate a self-supervised fine-tuning approach for T2I alignment based on RFMI that does not require auxiliary information other than the pre-trained model itself. Specifically, a fine-tuning set is constructed by selecting synthetic images generated from the pre-trained RF model and having high point-wise MI between images and prompts. Our experiments on MI estimation benchmarks demonstrate the validity of RFMI, and empirical fine-tuning on SD3.5-Medium confirms the effectiveness of RFMI for improving T2I alignment while maintaining image quality.
format Preprint
id arxiv_https___arxiv_org_abs_2503_14358
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RFMI: Estimating Mutual Information on Rectified Flow for Text-to-Image Alignment
Wang, Chao
Franzese, Giulio
Finamore, Alessandro
Michiardi, Pietro
Computer Vision and Pattern Recognition
Machine Learning
Rectified Flow (RF) models trained with a Flow matching framework have achieved state-of-the-art performance on Text-to-Image (T2I) conditional generation. Yet, multiple benchmarks show that synthetic images can still suffer from poor alignment with the prompt, i.e., images show wrong attribute binding, subject positioning, numeracy, etc. While the literature offers many methods to improve T2I alignment, they all consider only Diffusion Models, and require auxiliary datasets, scoring models, and linguistic analysis of the prompt. In this paper we aim to address these gaps. First, we introduce RFMI, a novel Mutual Information (MI) estimator for RF models that uses the pre-trained model itself for the MI estimation. Then, we investigate a self-supervised fine-tuning approach for T2I alignment based on RFMI that does not require auxiliary information other than the pre-trained model itself. Specifically, a fine-tuning set is constructed by selecting synthetic images generated from the pre-trained RF model and having high point-wise MI between images and prompts. Our experiments on MI estimation benchmarks demonstrate the validity of RFMI, and empirical fine-tuning on SD3.5-Medium confirms the effectiveness of RFMI for improving T2I alignment while maintaining image quality.
title RFMI: Estimating Mutual Information on Rectified Flow for Text-to-Image Alignment
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2503.14358