Large Multimodal Models for Low-Resource Languages: A Survey

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lupascu, Marian, Rogoz, Ana-Cristina, Stupariu, Mihai Sorin, Ionescu, Radu Tudor
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915767710121984
author Lupascu, Marian
Rogoz, Ana-Cristina
Stupariu, Mihai Sorin
Ionescu, Radu Tudor
author_facet Lupascu, Marian
Rogoz, Ana-Cristina
Stupariu, Mihai Sorin
Ionescu, Radu Tudor
contents In this survey, we systematically analyze techniques used to adapt large multimodal models (LMMs) for low-resource (LR) languages, examining approaches ranging from visual enhancement and data creation to cross-modal transfer and fusion strategies. Through a comprehensive analysis of 117 studies across 96 LR languages, we identify key patterns in how researchers tackle the challenges of limited data and computational resources. We categorize works into resource-oriented and method-oriented contributions, further dividing contributions into relevant sub-categories. We compare method-oriented contributions in terms of performance and efficiency, discussing benefits and limitations of representative studies. We find that visual information often serves as a crucial bridge for improving model performance in LR settings, though significant challenges remain in areas such as hallucination mitigation and computational efficiency. In summary, we provide researchers with a clear understanding of current approaches and remaining challenges in making LMMs more accessible to speakers of LR (understudied) languages. We complement our survey with an open-source repository available at: https://github.com/marianlupascu/LMM4LRL-Survey.
format Preprint
id arxiv_https___arxiv_org_abs_2502_05568
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Large Multimodal Models for Low-Resource Languages: A Survey
Lupascu, Marian
Rogoz, Ana-Cristina
Stupariu, Mihai Sorin
Ionescu, Radu Tudor
Computation and Language
Artificial Intelligence
Machine Learning
In this survey, we systematically analyze techniques used to adapt large multimodal models (LMMs) for low-resource (LR) languages, examining approaches ranging from visual enhancement and data creation to cross-modal transfer and fusion strategies. Through a comprehensive analysis of 117 studies across 96 LR languages, we identify key patterns in how researchers tackle the challenges of limited data and computational resources. We categorize works into resource-oriented and method-oriented contributions, further dividing contributions into relevant sub-categories. We compare method-oriented contributions in terms of performance and efficiency, discussing benefits and limitations of representative studies. We find that visual information often serves as a crucial bridge for improving model performance in LR settings, though significant challenges remain in areas such as hallucination mitigation and computational efficiency. In summary, we provide researchers with a clear understanding of current approaches and remaining challenges in making LMMs more accessible to speakers of LR (understudied) languages. We complement our survey with an open-source repository available at: https://github.com/marianlupascu/LMM4LRL-Survey.
title Large Multimodal Models for Low-Resource Languages: A Survey
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2502.05568