Text-centric Alignment for Multi-Modality Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Tsai, Yun-Da, Yen, Ting-Yu, Guo, Pei-Fu, Li, Zhe-Yan, Lin, Shou-De
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914803650396160
author Tsai, Yun-Da
Yen, Ting-Yu
Guo, Pei-Fu
Li, Zhe-Yan
Lin, Shou-De
author_facet Tsai, Yun-Da
Yen, Ting-Yu
Guo, Pei-Fu
Li, Zhe-Yan
Lin, Shou-De
contents This research paper addresses the challenge of modality mismatch in multimodal learning, where the modalities available during inference differ from those available at training. We propose the Text-centric Alignment for Multi-Modality Learning (TAMML) approach, an innovative method that utilizes Large Language Models (LLMs) with in-context learning and foundation models to enhance the generalizability of multimodal systems under these conditions. By leveraging the unique properties of text as a unified semantic space, TAMML demonstrates significant improvements in handling unseen, diverse, and unpredictable modality combinations. TAMML not only adapts to varying modalities but also maintains robust performance, showcasing the potential of foundation models in overcoming the limitations of traditional fixed-modality frameworks in embedding representations. This study contributes to the field by offering a flexible, effective solution for real-world applications where modality availability is dynamic and uncertain.
format Preprint
id arxiv_https___arxiv_org_abs_2402_08086
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Text-centric Alignment for Multi-Modality Learning
Tsai, Yun-Da
Yen, Ting-Yu
Guo, Pei-Fu
Li, Zhe-Yan
Lin, Shou-De
Machine Learning
Computation and Language
Computer Vision and Pattern Recognition
This research paper addresses the challenge of modality mismatch in multimodal learning, where the modalities available during inference differ from those available at training. We propose the Text-centric Alignment for Multi-Modality Learning (TAMML) approach, an innovative method that utilizes Large Language Models (LLMs) with in-context learning and foundation models to enhance the generalizability of multimodal systems under these conditions. By leveraging the unique properties of text as a unified semantic space, TAMML demonstrates significant improvements in handling unseen, diverse, and unpredictable modality combinations. TAMML not only adapts to varying modalities but also maintains robust performance, showcasing the potential of foundation models in overcoming the limitations of traditional fixed-modality frameworks in embedding representations. This study contributes to the field by offering a flexible, effective solution for real-world applications where modality availability is dynamic and uncertain.
title Text-centric Alignment for Multi-Modality Learning
topic Machine Learning
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2402.08086