Distributed LLMs and Multimodal Large Language Models: A Survey on Advances, Challenges, and Future Directions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Amini, Hadi, Mia, Md Jueal, Saadati, Yasaman, Imteaj, Ahmed, Nabavirazavi, Seyedsina, Thakker, Urmish, Hossain, Md Zarif, Fime, Awal Ahmed, Iyengar, S. S.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909546962747392
author Amini, Hadi
Mia, Md Jueal
Saadati, Yasaman
Imteaj, Ahmed
Nabavirazavi, Seyedsina
Thakker, Urmish
Hossain, Md Zarif
Fime, Awal Ahmed
Iyengar, S. S.
author_facet Amini, Hadi
Mia, Md Jueal
Saadati, Yasaman
Imteaj, Ahmed
Nabavirazavi, Seyedsina
Thakker, Urmish
Hossain, Md Zarif
Fime, Awal Ahmed
Iyengar, S. S.
contents Language models (LMs) are machine learning models designed to predict linguistic patterns by estimating the probability of word sequences based on large-scale datasets, such as text. LMs have a wide range of applications in natural language processing (NLP) tasks, including autocomplete and machine translation. Although larger datasets typically enhance LM performance, scalability remains a challenge due to constraints in computational power and resources. Distributed computing strategies offer essential solutions for improving scalability and managing the growing computational demand. Further, the use of sensitive datasets in training and deployment raises significant privacy concerns. Recent research has focused on developing decentralized techniques to enable distributed training and inference while utilizing diverse computational resources and enabling edge AI. This paper presents a survey on distributed solutions for various LMs, including large language models (LLMs), vision language models (VLMs), multimodal LLMs (MLLMs), and small language models (SLMs). While LLMs focus on processing and generating text, MLLMs are designed to handle multiple modalities of data (e.g., text, images, and audio) and to integrate them for broader applications. To this end, this paper reviews key advancements across the MLLM pipeline, including distributed training, inference, fine-tuning, and deployment, while also identifying the contributions, limitations, and future areas of improvement. Further, it categorizes the literature based on six primary focus areas of decentralization. Our analysis describes gaps in current methodologies for enabling distributed solutions for LMs and outline future research directions, emphasizing the need for novel solutions to enhance the robustness and applicability of distributed LMs.
format Preprint
id arxiv_https___arxiv_org_abs_2503_16585
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Distributed LLMs and Multimodal Large Language Models: A Survey on Advances, Challenges, and Future Directions
Amini, Hadi
Mia, Md Jueal
Saadati, Yasaman
Imteaj, Ahmed
Nabavirazavi, Seyedsina
Thakker, Urmish
Hossain, Md Zarif
Fime, Awal Ahmed
Iyengar, S. S.
Computation and Language
Computer Vision and Pattern Recognition
Distributed, Parallel, and Cluster Computing
Machine Learning
Language models (LMs) are machine learning models designed to predict linguistic patterns by estimating the probability of word sequences based on large-scale datasets, such as text. LMs have a wide range of applications in natural language processing (NLP) tasks, including autocomplete and machine translation. Although larger datasets typically enhance LM performance, scalability remains a challenge due to constraints in computational power and resources. Distributed computing strategies offer essential solutions for improving scalability and managing the growing computational demand. Further, the use of sensitive datasets in training and deployment raises significant privacy concerns. Recent research has focused on developing decentralized techniques to enable distributed training and inference while utilizing diverse computational resources and enabling edge AI. This paper presents a survey on distributed solutions for various LMs, including large language models (LLMs), vision language models (VLMs), multimodal LLMs (MLLMs), and small language models (SLMs). While LLMs focus on processing and generating text, MLLMs are designed to handle multiple modalities of data (e.g., text, images, and audio) and to integrate them for broader applications. To this end, this paper reviews key advancements across the MLLM pipeline, including distributed training, inference, fine-tuning, and deployment, while also identifying the contributions, limitations, and future areas of improvement. Further, it categorizes the literature based on six primary focus areas of decentralization. Our analysis describes gaps in current methodologies for enabling distributed solutions for LMs and outline future research directions, emphasizing the need for novel solutions to enhance the robustness and applicability of distributed LMs.
title Distributed LLMs and Multimodal Large Language Models: A Survey on Advances, Challenges, and Future Directions
topic Computation and Language
Computer Vision and Pattern Recognition
Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2503.16585