ModalChorus: Visual Probing and Alignment of Multi-modal Embeddings via Modal Fusion Map

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ye, Yilin, Xiao, Shishi, Zeng, Xingchen, Zeng, Wei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916455651475456
author Ye, Yilin
Xiao, Shishi
Zeng, Xingchen
Zeng, Wei
author_facet Ye, Yilin
Xiao, Shishi
Zeng, Xingchen
Zeng, Wei
contents Multi-modal embeddings form the foundation for vision-language models, such as CLIP embeddings, the most widely used text-image embeddings. However, these embeddings are vulnerable to subtle misalignment of cross-modal features, resulting in decreased model performance and diminished generalization. To address this problem, we design ModalChorus, an interactive system for visual probing and alignment of multi-modal embeddings. ModalChorus primarily offers a two-stage process: 1) embedding probing with Modal Fusion Map (MFM), a novel parametric dimensionality reduction method that integrates both metric and nonmetric objectives to enhance modality fusion; and 2) embedding alignment that allows users to interactively articulate intentions for both point-set and set-set alignments. Quantitative and qualitative comparisons for CLIP embeddings with existing dimensionality reduction (e.g., t-SNE and MDS) and data fusion (e.g., data context map) methods demonstrate the advantages of MFM in showcasing cross-modal features over common vision-language datasets. Case studies reveal that ModalChorus can facilitate intuitive discovery of misalignment and efficient re-alignment in scenarios ranging from zero-shot classification to cross-modal retrieval and generation.
format Preprint
id arxiv_https___arxiv_org_abs_2407_12315
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ModalChorus: Visual Probing and Alignment of Multi-modal Embeddings via Modal Fusion Map
Ye, Yilin
Xiao, Shishi
Zeng, Xingchen
Zeng, Wei
Computer Vision and Pattern Recognition
Artificial Intelligence
Human-Computer Interaction
Information Retrieval
Multi-modal embeddings form the foundation for vision-language models, such as CLIP embeddings, the most widely used text-image embeddings. However, these embeddings are vulnerable to subtle misalignment of cross-modal features, resulting in decreased model performance and diminished generalization. To address this problem, we design ModalChorus, an interactive system for visual probing and alignment of multi-modal embeddings. ModalChorus primarily offers a two-stage process: 1) embedding probing with Modal Fusion Map (MFM), a novel parametric dimensionality reduction method that integrates both metric and nonmetric objectives to enhance modality fusion; and 2) embedding alignment that allows users to interactively articulate intentions for both point-set and set-set alignments. Quantitative and qualitative comparisons for CLIP embeddings with existing dimensionality reduction (e.g., t-SNE and MDS) and data fusion (e.g., data context map) methods demonstrate the advantages of MFM in showcasing cross-modal features over common vision-language datasets. Case studies reveal that ModalChorus can facilitate intuitive discovery of misalignment and efficient re-alignment in scenarios ranging from zero-shot classification to cross-modal retrieval and generation.
title ModalChorus: Visual Probing and Alignment of Multi-modal Embeddings via Modal Fusion Map
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Human-Computer Interaction
Information Retrieval
url https://arxiv.org/abs/2407.12315