If you can describe it, they can see it: Cross-Modal Learning of Visual Concepts from Textual Descriptions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Barbano, Carlo Alberto, Molinaro, Luca, Ciranni, Massimiliano, Aiello, Emanuele, Pastore, Vito Paolo, Grangetto, Marco
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909965712621568
author Barbano, Carlo Alberto
Molinaro, Luca
Ciranni, Massimiliano
Aiello, Emanuele
Pastore, Vito Paolo
Grangetto, Marco
author_facet Barbano, Carlo Alberto
Molinaro, Luca
Ciranni, Massimiliano
Aiello, Emanuele
Pastore, Vito Paolo
Grangetto, Marco
contents Humans can visualize new and unknown concepts from their natural language description, based on their experience and previous knowledge. Insipired by this, we present a way to extend this ability to Vision-Language Models (VLMs), teaching them novel concepts by only using a textual description. We refer to this approach as Knowledge Transfer (KT). Our hypothesis is that the knowledge of a pre-trained VLM can be re-used to represent previously unknown concepts. Provided with a textual description of the novel concept, KT works by aligning relevant features of the visual encoder, obtained through model inversion, to its text representation. Differently from approaches relying on visual examples or external generative models, KT transfers knowledge within the same VLM by injecting visual knowledge directly from the text. Through an extensive evaluation on several VLM tasks, including classification, segmentation, image-text retrieval, and captioning, we show that: 1) KT can efficiently introduce new visual concepts from a single textual description; 2) the same principle can be used to refine the representation of existing concepts; and 3) KT significantly improves the performance of zero-shot VLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2411_15611
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle If you can describe it, they can see it: Cross-Modal Learning of Visual Concepts from Textual Descriptions
Barbano, Carlo Alberto
Molinaro, Luca
Ciranni, Massimiliano
Aiello, Emanuele
Pastore, Vito Paolo
Grangetto, Marco
Computer Vision and Pattern Recognition
68T45 (Primary) 68T50 (Secondary)
I.2.6
Humans can visualize new and unknown concepts from their natural language description, based on their experience and previous knowledge. Insipired by this, we present a way to extend this ability to Vision-Language Models (VLMs), teaching them novel concepts by only using a textual description. We refer to this approach as Knowledge Transfer (KT). Our hypothesis is that the knowledge of a pre-trained VLM can be re-used to represent previously unknown concepts. Provided with a textual description of the novel concept, KT works by aligning relevant features of the visual encoder, obtained through model inversion, to its text representation. Differently from approaches relying on visual examples or external generative models, KT transfers knowledge within the same VLM by injecting visual knowledge directly from the text. Through an extensive evaluation on several VLM tasks, including classification, segmentation, image-text retrieval, and captioning, we show that: 1) KT can efficiently introduce new visual concepts from a single textual description; 2) the same principle can be used to refine the representation of existing concepts; and 3) KT significantly improves the performance of zero-shot VLMs.
title If you can describe it, they can see it: Cross-Modal Learning of Visual Concepts from Textual Descriptions
topic Computer Vision and Pattern Recognition
68T45 (Primary) 68T50 (Secondary)
I.2.6
url https://arxiv.org/abs/2411.15611