Beyond Monolingual Assumptions: A Survey of Code-Switched NLP in the Era of Large Language Models across Modalities

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Sheth, Rajvee, Sinha, Samridhi Raj, Patil, Mahavir, Beniwal, Himanshu, Singh, Mayank
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914493931454464
author Sheth, Rajvee
Sinha, Samridhi Raj
Patil, Mahavir
Beniwal, Himanshu
Singh, Mayank
author_facet Sheth, Rajvee
Sinha, Samridhi Raj
Patil, Mahavir
Beniwal, Himanshu
Singh, Mayank
contents Amidst the rapid advances of large language models (LLMs), most LLMs still struggle with mixed-language inputs, limited Codeswitching (CSW) datasets, and evaluation biases, which hinder their deployment in multilingual societies. This survey provides the first comprehensive analysis of CSW-aware LLM research, reviewing 327 studies spanning five research areas, 15+ NLP tasks, 30+ datasets, and 80+ languages. We categorize recent advances by architecture, training strategy, and evaluation methodology, outlining how LLMs have reshaped CSW modeling and identifying the challenges that persist. The paper concludes with a roadmap that emphasizes the need for inclusive datasets, fair evaluation, and linguistically grounded models to achieve truly multilingual capabilities https://github.com/lingo-iitgn/awesome-code-mixing/.
format Preprint
id arxiv_https___arxiv_org_abs_2510_07037
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Monolingual Assumptions: A Survey of Code-Switched NLP in the Era of Large Language Models across Modalities
Sheth, Rajvee
Sinha, Samridhi Raj
Patil, Mahavir
Beniwal, Himanshu
Singh, Mayank
Computation and Language
Amidst the rapid advances of large language models (LLMs), most LLMs still struggle with mixed-language inputs, limited Codeswitching (CSW) datasets, and evaluation biases, which hinder their deployment in multilingual societies. This survey provides the first comprehensive analysis of CSW-aware LLM research, reviewing 327 studies spanning five research areas, 15+ NLP tasks, 30+ datasets, and 80+ languages. We categorize recent advances by architecture, training strategy, and evaluation methodology, outlining how LLMs have reshaped CSW modeling and identifying the challenges that persist. The paper concludes with a roadmap that emphasizes the need for inclusive datasets, fair evaluation, and linguistically grounded models to achieve truly multilingual capabilities https://github.com/lingo-iitgn/awesome-code-mixing/.
title Beyond Monolingual Assumptions: A Survey of Code-Switched NLP in the Era of Large Language Models across Modalities
topic Computation and Language
url https://arxiv.org/abs/2510.07037