Discffusion: Discriminative Diffusion Models as Few-shot Vision and Language Learners
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866929326882029568 |
|---|---|
| author | He, Xuehai Feng, Weixi Fu, Tsu-Jui Jampani, Varun Akula, Arjun Narayana, Pradyumna Basu, Sugato Wang, William Yang Wang, Xin Eric |
| author_facet | He, Xuehai Feng, Weixi Fu, Tsu-Jui Jampani, Varun Akula, Arjun Narayana, Pradyumna Basu, Sugato Wang, William Yang Wang, Xin Eric |
| contents | Diffusion models, such as Stable Diffusion, have shown incredible performance on text-to-image generation. Since text-to-image generation often requires models to generate visual concepts with fine-grained details and attributes specified in text prompts, can we leverage the powerful representations learned by pre-trained diffusion models for discriminative tasks such as image-text matching? To answer this question, we propose a novel approach, Discriminative Stable Diffusion (DSD), which turns pre-trained text-to-image diffusion models into few-shot discriminative learners. Our approach mainly uses the cross-attention score of a Stable Diffusion model to capture the mutual influence between visual and textual information and fine-tune the model via efficient attention-based prompt learning to perform image-text matching. By comparing DSD with state-of-the-art methods on several benchmark datasets, we demonstrate the potential of using pre-trained diffusion models for discriminative tasks with superior results on few-shot image-text matching. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2305_10722 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Discffusion: Discriminative Diffusion Models as Few-shot Vision and Language Learners He, Xuehai Feng, Weixi Fu, Tsu-Jui Jampani, Varun Akula, Arjun Narayana, Pradyumna Basu, Sugato Wang, William Yang Wang, Xin Eric Computer Vision and Pattern Recognition Diffusion models, such as Stable Diffusion, have shown incredible performance on text-to-image generation. Since text-to-image generation often requires models to generate visual concepts with fine-grained details and attributes specified in text prompts, can we leverage the powerful representations learned by pre-trained diffusion models for discriminative tasks such as image-text matching? To answer this question, we propose a novel approach, Discriminative Stable Diffusion (DSD), which turns pre-trained text-to-image diffusion models into few-shot discriminative learners. Our approach mainly uses the cross-attention score of a Stable Diffusion model to capture the mutual influence between visual and textual information and fine-tune the model via efficient attention-based prompt learning to perform image-text matching. By comparing DSD with state-of-the-art methods on several benchmark datasets, we demonstrate the potential of using pre-trained diffusion models for discriminative tasks with superior results on few-shot image-text matching. |
| title | Discffusion: Discriminative Diffusion Models as Few-shot Vision and Language Learners |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2305.10722 |