CACARA: Cross-Modal Alignment Leveraging a Text-Centric Approach for Cost-Effective Multimodal and Multilingual Learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Moreira, Diego A. B., Ferreira, Alef I., Silva, Jhessica, Santos, Gabriel O. dos, Bonil, Gustavo, Gondim, João, Santos, Marina dos, Maia, Helena, Hashiguti, Simone, da Silva, Nádia, Scarton, Carolina, Pedrini, Helio, Avila, Sandra
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911294395777024
author Moreira, Diego A. B.
Ferreira, Alef I.
Silva, Jhessica
Santos, Gabriel O. dos
Bonil, Gustavo
Gondim, João
Santos, Marina dos
Maia, Helena
Hashiguti, Simone
da Silva, Nádia
Scarton, Carolina
Pedrini, Helio
Avila, Sandra
author_facet Moreira, Diego A. B.
Ferreira, Alef I.
Silva, Jhessica
Santos, Gabriel O. dos
Bonil, Gustavo
Gondim, João
Santos, Marina dos
Maia, Helena
Hashiguti, Simone
da Silva, Nádia
Scarton, Carolina
Pedrini, Helio
Avila, Sandra
contents As deep learning models evolve, new applications and challenges are rapidly emerging. Tasks that once relied on a single modality, such as text, images, or audio, are now enriched by seamless interactions between multimodal data. These connections bridge information gaps: an image can visually materialize a text, while audio can add context to an image. Researchers have developed numerous multimodal models, but most rely on resource-intensive training across multiple modalities. Similarly, extending these models to new languages often follows the same resource-heavy training strategy. In this work, we propose a multimodal and multilingual architecture, CACARA, trained through emergent alignment learning, enabling the seamless integration of new modalities into an existing bimodal/multimodal model without requiring full retraining. This work breaks new ground by demonstrating that this emergent alignment paradigm can unlock multilingual capabilities from monolingual training. By fine-tuning the newly incorporated modality only on data aligned with the English language, our model develops support for over 100 languages without explicit multilingual pretraining or tuning of the text encoder. Such emergent multimodal and multilingual properties are gained efficiently, preserving previously learned knowledge at a training cost comparable to that of a monolingual model. Our strategy achieves up to a 14.24 percentage points improvement in R@1 audio-to-text retrieval, outperforming state-of-the-art multimodal models -- all without the heavy computational cost of retraining across every modality and language.
format Preprint
id arxiv_https___arxiv_org_abs_2512_00496
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CACARA: Cross-Modal Alignment Leveraging a Text-Centric Approach for Cost-Effective Multimodal and Multilingual Learning
Moreira, Diego A. B.
Ferreira, Alef I.
Silva, Jhessica
Santos, Gabriel O. dos
Bonil, Gustavo
Gondim, João
Santos, Marina dos
Maia, Helena
Hashiguti, Simone
da Silva, Nádia
Scarton, Carolina
Pedrini, Helio
Avila, Sandra
Computation and Language
Artificial Intelligence
As deep learning models evolve, new applications and challenges are rapidly emerging. Tasks that once relied on a single modality, such as text, images, or audio, are now enriched by seamless interactions between multimodal data. These connections bridge information gaps: an image can visually materialize a text, while audio can add context to an image. Researchers have developed numerous multimodal models, but most rely on resource-intensive training across multiple modalities. Similarly, extending these models to new languages often follows the same resource-heavy training strategy. In this work, we propose a multimodal and multilingual architecture, CACARA, trained through emergent alignment learning, enabling the seamless integration of new modalities into an existing bimodal/multimodal model without requiring full retraining. This work breaks new ground by demonstrating that this emergent alignment paradigm can unlock multilingual capabilities from monolingual training. By fine-tuning the newly incorporated modality only on data aligned with the English language, our model develops support for over 100 languages without explicit multilingual pretraining or tuning of the text encoder. Such emergent multimodal and multilingual properties are gained efficiently, preserving previously learned knowledge at a training cost comparable to that of a monolingual model. Our strategy achieves up to a 14.24 percentage points improvement in R@1 audio-to-text retrieval, outperforming state-of-the-art multimodal models -- all without the heavy computational cost of retraining across every modality and language.
title CACARA: Cross-Modal Alignment Leveraging a Text-Centric Approach for Cost-Effective Multimodal and Multilingual Learning
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2512.00496