CSA: Data-efficient Mapping of Unimodal Features to Multimodal Features

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Po-han, Chinchali, Sandeep P., Topcu, Ufuk
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909536987643904
author Li, Po-han
Chinchali, Sandeep P.
Topcu, Ufuk
author_facet Li, Po-han
Chinchali, Sandeep P.
Topcu, Ufuk
contents Multimodal encoders like CLIP excel in tasks such as zero-shot image classification and cross-modal retrieval. However, they require excessive training data. We propose canonical similarity analysis (CSA), which uses two unimodal encoders to replicate multimodal encoders using limited data. CSA maps unimodal features into a multimodal space, using a new similarity score to retain only the multimodal information. CSA only involves the inference of unimodal encoders and a cubic-complexity matrix decomposition, eliminating the need for extensive GPU-based model training. Experiments show that CSA outperforms CLIP while requiring $50,000\times$ fewer multimodal data pairs to bridge the modalities given pre-trained unimodal encoders on ImageNet classification and misinformative news caption detection. CSA surpasses the state-of-the-art method to map unimodal features to multimodal features. We also demonstrate the ability of CSA with modalities beyond image and text, paving the way for future modality pairs with limited paired multimodal data but abundant unpaired unimodal data, such as lidar and text.
format Preprint
id arxiv_https___arxiv_org_abs_2410_07610
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CSA: Data-efficient Mapping of Unimodal Features to Multimodal Features
Li, Po-han
Chinchali, Sandeep P.
Topcu, Ufuk
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Information Retrieval
Multimodal encoders like CLIP excel in tasks such as zero-shot image classification and cross-modal retrieval. However, they require excessive training data. We propose canonical similarity analysis (CSA), which uses two unimodal encoders to replicate multimodal encoders using limited data. CSA maps unimodal features into a multimodal space, using a new similarity score to retain only the multimodal information. CSA only involves the inference of unimodal encoders and a cubic-complexity matrix decomposition, eliminating the need for extensive GPU-based model training. Experiments show that CSA outperforms CLIP while requiring $50,000\times$ fewer multimodal data pairs to bridge the modalities given pre-trained unimodal encoders on ImageNet classification and misinformative news caption detection. CSA surpasses the state-of-the-art method to map unimodal features to multimodal features. We also demonstrate the ability of CSA with modalities beyond image and text, paving the way for future modality pairs with limited paired multimodal data but abundant unpaired unimodal data, such as lidar and text.
title CSA: Data-efficient Mapping of Unimodal Features to Multimodal Features
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Information Retrieval
url https://arxiv.org/abs/2410.07610