Any2Any: Incomplete Multimodal Retrieval with Conformal Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Po-han, Yang, Yunhao, Omama, Mohammad, Chinchali, Sandeep, Topcu, Ufuk
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929604186341376
author Li, Po-han
Yang, Yunhao
Omama, Mohammad
Chinchali, Sandeep
Topcu, Ufuk
author_facet Li, Po-han
Yang, Yunhao
Omama, Mohammad
Chinchali, Sandeep
Topcu, Ufuk
contents Autonomous agents perceive and interpret their surroundings by integrating multimodal inputs, such as vision, audio, and LiDAR. These perceptual modalities support retrieval tasks, such as place recognition in robotics. However, current multimodal retrieval systems encounter difficulties when parts of the data are missing due to sensor failures or inaccessibility, such as silent videos or LiDAR scans lacking RGB information. We propose Any2Any-a novel retrieval framework that addresses scenarios where both query and reference instances have incomplete modalities. Unlike previous methods limited to the imputation of two modalities, Any2Any handles any number of modalities without training generative models. It calculates pairwise similarities with cross-modal encoders and employs a two-stage calibration process with conformal prediction to align the similarities. Any2Any enables effective retrieval across multimodal datasets, e.g., text-LiDAR and text-time series. It achieves a Recall@5 of 35% on the KITTI dataset, which is on par with baseline models with complete modalities.
format Preprint
id arxiv_https___arxiv_org_abs_2411_10513
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Any2Any: Incomplete Multimodal Retrieval with Conformal Prediction
Li, Po-han
Yang, Yunhao
Omama, Mohammad
Chinchali, Sandeep
Topcu, Ufuk
Computer Vision and Pattern Recognition
Information Retrieval
Multimedia
Autonomous agents perceive and interpret their surroundings by integrating multimodal inputs, such as vision, audio, and LiDAR. These perceptual modalities support retrieval tasks, such as place recognition in robotics. However, current multimodal retrieval systems encounter difficulties when parts of the data are missing due to sensor failures or inaccessibility, such as silent videos or LiDAR scans lacking RGB information. We propose Any2Any-a novel retrieval framework that addresses scenarios where both query and reference instances have incomplete modalities. Unlike previous methods limited to the imputation of two modalities, Any2Any handles any number of modalities without training generative models. It calculates pairwise similarities with cross-modal encoders and employs a two-stage calibration process with conformal prediction to align the similarities. Any2Any enables effective retrieval across multimodal datasets, e.g., text-LiDAR and text-time series. It achieves a Recall@5 of 35% on the KITTI dataset, which is on par with baseline models with complete modalities.
title Any2Any: Incomplete Multimodal Retrieval with Conformal Prediction
topic Computer Vision and Pattern Recognition
Information Retrieval
Multimedia
url https://arxiv.org/abs/2411.10513