Discffusion: Discriminative Diffusion Models as Few-shot Vision and Language Learners

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: He, Xuehai, Feng, Weixi, Fu, Tsu-Jui, Jampani, Varun, Akula, Arjun, Narayana, Pradyumna, Basu, Sugato, Wang, William Yang, Wang, Xin Eric
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929326882029568
author He, Xuehai
Feng, Weixi
Fu, Tsu-Jui
Jampani, Varun
Akula, Arjun
Narayana, Pradyumna
Basu, Sugato
Wang, William Yang
Wang, Xin Eric
author_facet He, Xuehai
Feng, Weixi
Fu, Tsu-Jui
Jampani, Varun
Akula, Arjun
Narayana, Pradyumna
Basu, Sugato
Wang, William Yang
Wang, Xin Eric
contents Diffusion models, such as Stable Diffusion, have shown incredible performance on text-to-image generation. Since text-to-image generation often requires models to generate visual concepts with fine-grained details and attributes specified in text prompts, can we leverage the powerful representations learned by pre-trained diffusion models for discriminative tasks such as image-text matching? To answer this question, we propose a novel approach, Discriminative Stable Diffusion (DSD), which turns pre-trained text-to-image diffusion models into few-shot discriminative learners. Our approach mainly uses the cross-attention score of a Stable Diffusion model to capture the mutual influence between visual and textual information and fine-tune the model via efficient attention-based prompt learning to perform image-text matching. By comparing DSD with state-of-the-art methods on several benchmark datasets, we demonstrate the potential of using pre-trained diffusion models for discriminative tasks with superior results on few-shot image-text matching.
format Preprint
id arxiv_https___arxiv_org_abs_2305_10722
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Discffusion: Discriminative Diffusion Models as Few-shot Vision and Language Learners
He, Xuehai
Feng, Weixi
Fu, Tsu-Jui
Jampani, Varun
Akula, Arjun
Narayana, Pradyumna
Basu, Sugato
Wang, William Yang
Wang, Xin Eric
Computer Vision and Pattern Recognition
Diffusion models, such as Stable Diffusion, have shown incredible performance on text-to-image generation. Since text-to-image generation often requires models to generate visual concepts with fine-grained details and attributes specified in text prompts, can we leverage the powerful representations learned by pre-trained diffusion models for discriminative tasks such as image-text matching? To answer this question, we propose a novel approach, Discriminative Stable Diffusion (DSD), which turns pre-trained text-to-image diffusion models into few-shot discriminative learners. Our approach mainly uses the cross-attention score of a Stable Diffusion model to capture the mutual influence between visual and textual information and fine-tune the model via efficient attention-based prompt learning to perform image-text matching. By comparing DSD with state-of-the-art methods on several benchmark datasets, we demonstrate the potential of using pre-trained diffusion models for discriminative tasks with superior results on few-shot image-text matching.
title Discffusion: Discriminative Diffusion Models as Few-shot Vision and Language Learners
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2305.10722