Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Agarwal, Vatsal, Gwilliam, Matthew, Kohavi, Gefen, Verma, Eshan, Ulbricht, Daniel, Shrivastava, Abhinav
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918087727513600
author Agarwal, Vatsal
Gwilliam, Matthew
Kohavi, Gefen
Verma, Eshan
Ulbricht, Daniel
Shrivastava, Abhinav
author_facet Agarwal, Vatsal
Gwilliam, Matthew
Kohavi, Gefen
Verma, Eshan
Ulbricht, Daniel
Shrivastava, Abhinav
contents Recent advances in multimodal large language models (MLLMs) have enabled image-based question-answering capabilities. However, a key limitation is the use of CLIP as the visual encoder; while it can capture coarse global information, it often can miss fine-grained details that are relevant to the input query. To address these shortcomings, this work studies whether pre-trained text-to-image diffusion models can serve as instruction-aware visual encoders. Through an analysis of their internal representations, we find diffusion features are both rich in semantics and can encode strong image-text alignment. Moreover, we find that we can leverage text conditioning to focus the model on regions relevant to the input question. We then investigate how to align these features with large language models and uncover a leakage phenomenon, where the LLM can inadvertently recover information from the original diffusion prompt. We analyze the causes of this leakage and propose a mitigation strategy. Based on these insights, we explore a simple fusion strategy that utilizes both CLIP and conditional diffusion features. We evaluate our approach on both general VQA and specialized MLLM benchmarks, demonstrating the promise of diffusion models for visual understanding, particularly in vision-centric tasks that require spatial and compositional reasoning. Our project page can be found https://vatsalag99.github.io/mustafar/.
format Preprint
id arxiv_https___arxiv_org_abs_2507_07106
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor
Agarwal, Vatsal
Gwilliam, Matthew
Kohavi, Gefen
Verma, Eshan
Ulbricht, Daniel
Shrivastava, Abhinav
Computer Vision and Pattern Recognition
Machine Learning
Recent advances in multimodal large language models (MLLMs) have enabled image-based question-answering capabilities. However, a key limitation is the use of CLIP as the visual encoder; while it can capture coarse global information, it often can miss fine-grained details that are relevant to the input query. To address these shortcomings, this work studies whether pre-trained text-to-image diffusion models can serve as instruction-aware visual encoders. Through an analysis of their internal representations, we find diffusion features are both rich in semantics and can encode strong image-text alignment. Moreover, we find that we can leverage text conditioning to focus the model on regions relevant to the input question. We then investigate how to align these features with large language models and uncover a leakage phenomenon, where the LLM can inadvertently recover information from the original diffusion prompt. We analyze the causes of this leakage and propose a mitigation strategy. Based on these insights, we explore a simple fusion strategy that utilizes both CLIP and conditional diffusion features. We evaluate our approach on both general VQA and specialized MLLM benchmarks, demonstrating the promise of diffusion models for visual understanding, particularly in vision-centric tasks that require spatial and compositional reasoning. Our project page can be found https://vatsalag99.github.io/mustafar/.
title Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2507.07106