If you can describe it, they can see it: Cross-Modal Learning of Visual Concepts from Textual Descriptions
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909965712621568 |
|---|---|
| author | Barbano, Carlo Alberto Molinaro, Luca Ciranni, Massimiliano Aiello, Emanuele Pastore, Vito Paolo Grangetto, Marco |
| author_facet | Barbano, Carlo Alberto Molinaro, Luca Ciranni, Massimiliano Aiello, Emanuele Pastore, Vito Paolo Grangetto, Marco |
| contents | Humans can visualize new and unknown concepts from their natural language description, based on their experience and previous knowledge. Insipired by this, we present a way to extend this ability to Vision-Language Models (VLMs), teaching them novel concepts by only using a textual description. We refer to this approach as Knowledge Transfer (KT). Our hypothesis is that the knowledge of a pre-trained VLM can be re-used to represent previously unknown concepts. Provided with a textual description of the novel concept, KT works by aligning relevant features of the visual encoder, obtained through model inversion, to its text representation. Differently from approaches relying on visual examples or external generative models, KT transfers knowledge within the same VLM by injecting visual knowledge directly from the text. Through an extensive evaluation on several VLM tasks, including classification, segmentation, image-text retrieval, and captioning, we show that: 1) KT can efficiently introduce new visual concepts from a single textual description; 2) the same principle can be used to refine the representation of existing concepts; and 3) KT significantly improves the performance of zero-shot VLMs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_15611 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | If you can describe it, they can see it: Cross-Modal Learning of Visual Concepts from Textual Descriptions Barbano, Carlo Alberto Molinaro, Luca Ciranni, Massimiliano Aiello, Emanuele Pastore, Vito Paolo Grangetto, Marco Computer Vision and Pattern Recognition 68T45 (Primary) 68T50 (Secondary) I.2.6 Humans can visualize new and unknown concepts from their natural language description, based on their experience and previous knowledge. Insipired by this, we present a way to extend this ability to Vision-Language Models (VLMs), teaching them novel concepts by only using a textual description. We refer to this approach as Knowledge Transfer (KT). Our hypothesis is that the knowledge of a pre-trained VLM can be re-used to represent previously unknown concepts. Provided with a textual description of the novel concept, KT works by aligning relevant features of the visual encoder, obtained through model inversion, to its text representation. Differently from approaches relying on visual examples or external generative models, KT transfers knowledge within the same VLM by injecting visual knowledge directly from the text. Through an extensive evaluation on several VLM tasks, including classification, segmentation, image-text retrieval, and captioning, we show that: 1) KT can efficiently introduce new visual concepts from a single textual description; 2) the same principle can be used to refine the representation of existing concepts; and 3) KT significantly improves the performance of zero-shot VLMs. |
| title | If you can describe it, they can see it: Cross-Modal Learning of Visual Concepts from Textual Descriptions |
| topic | Computer Vision and Pattern Recognition 68T45 (Primary) 68T50 (Secondary) I.2.6 |
| url | https://arxiv.org/abs/2411.15611 |