Lost in the Mix: Evaluating LLM Understanding of Code-Switched Text
Fuente:
arXiv
Saved in:
| Main Authors: | Mohamed, Amr, Zhang, Yang, Vazirgiannis, Michalis, Shang, Guokan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fast-Decoding Diffusion Language Models via Progress-Aware Confidence Schedules
by: Mohamed, Amr, et al.
Published: (2025)
by: Mohamed, Amr, et al.
Published: (2025)
LLM as a Broken Telephone: Iterative Generation Distorts Information
by: Mohamed, Amr, et al.
Published: (2025)
by: Mohamed, Amr, et al.
Published: (2025)
Beyond Random Sampling: Efficient Language Model Pretraining via Curriculum Learning
by: Zhang, Yang, et al.
Published: (2025)
by: Zhang, Yang, et al.
Published: (2025)
The Curious Decline of Linguistic Diversity: Training Language Models on Synthetic Text
by: Guo, Yanzhu, et al.
Published: (2023)
by: Guo, Yanzhu, et al.
Published: (2023)
Markovian Generation Chains in Large Language Models
by: Geng, Mingmeng, et al.
Published: (2026)
by: Geng, Mingmeng, et al.
Published: (2026)
Leveraging Discourse Structure for Extractive Meeting Summarization
by: Rennard, Virgile, et al.
Published: (2024)
by: Rennard, Virgile, et al.
Published: (2024)
GreekMMLU: A Native-Sourced Multitask Benchmark for Evaluating Language Models in Greek
by: Zhang, Yang, et al.
Published: (2026)
by: Zhang, Yang, et al.
Published: (2026)
Nile-Chat: Egyptian Language Models for Arabic and Latin Scripts
by: Shang, Guokan, et al.
Published: (2025)
by: Shang, Guokan, et al.
Published: (2025)
Atlas-Chat: Adapting Large Language Models for Low-Resource Moroccan Arabic Dialect
by: Shang, Guokan, et al.
Published: (2024)
by: Shang, Guokan, et al.
Published: (2024)
Graph Linearization Methods for Reasoning on Graphs with Large Language Models
by: Xypolopoulos, Christos, et al.
Published: (2024)
by: Xypolopoulos, Christos, et al.
Published: (2024)
Prot2Text: Multimodal Protein's Function Generation with GNNs and Transformers
by: Abdine, Hadi, et al.
Published: (2023)
by: Abdine, Hadi, et al.
Published: (2023)
Bias in the Mirror: Are LLMs opinions robust to their own adversarial attacks ?
by: Rennard, Virgile, et al.
Published: (2024)
by: Rennard, Virgile, et al.
Published: (2024)
Code-Switching Red-Teaming: LLM Evaluation for Safety and Multilingual Understanding
by: Yoo, Haneul, et al.
Published: (2024)
by: Yoo, Haneul, et al.
Published: (2024)
Benchmarking Linguistic Diversity of Large Language Models
by: Guo, Yanzhu, et al.
Published: (2024)
by: Guo, Yanzhu, et al.
Published: (2024)
Graph Neural Networks on Discriminative Graphs of Words
by: Abbahaddou, Yassine, et al.
Published: (2024)
by: Abbahaddou, Yassine, et al.
Published: (2024)
CARTE: A Benchmark for Mapping Language Model Knowledge Across France
by: Carneiro, Sarah Almeida, et al.
Published: (2026)
by: Carneiro, Sarah Almeida, et al.
Published: (2026)
Lost in Speech: Benchmarking, Evaluation, and Parsing of Spoken Code-Switching Beyond Standard UD Assumptions
by: Tyagi, Nemika, et al.
Published: (2026)
by: Tyagi, Nemika, et al.
Published: (2026)
Word Sense Induction with Hierarchical Clustering and Mutual Information Maximization
by: Abdine, Hadi, et al.
Published: (2022)
by: Abdine, Hadi, et al.
Published: (2022)
Code-Mixed Probes Show How Pre-Trained Models Generalise On Code-Switched Text
by: De Leon, Frances A. Laureano, et al.
Published: (2024)
by: De Leon, Frances A. Laureano, et al.
Published: (2024)
Shorter but not Worse: Frugal Reasoning via Easy Samples as Length Regularizers in Math RLVR
by: Bounhar, Abdelaziz, et al.
Published: (2025)
by: Bounhar, Abdelaziz, et al.
Published: (2025)
LLM-based Code-Switched Text Generation for Grammatical Error Correction
by: Potter, Tom, et al.
Published: (2024)
by: Potter, Tom, et al.
Published: (2024)
Can Large Language Models Understand, Reason About, and Generate Code-Switched Text?
by: Winata, Genta Indra, et al.
Published: (2026)
by: Winata, Genta Indra, et al.
Published: (2026)
An Evaluation of Large Language Models on Text Summarization Tasks Using Prompt Engineering Techniques
by: Aly, Walid Mohamed, et al.
Published: (2025)
by: Aly, Walid Mohamed, et al.
Published: (2025)
CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages
by: Yang, Yilun, et al.
Published: (2025)
by: Yang, Yilun, et al.
Published: (2025)
Can Code-Switched Texts Activate a Knowledge Switch in LLMs? A Case Study on English-Korean Code-Switching
by: Kim, Seoyeon, et al.
Published: (2024)
by: Kim, Seoyeon, et al.
Published: (2024)
A Greek Government Decisions Dataset for Public-Sector Analysis and Insight
by: Antoniou, Giorgos, et al.
Published: (2025)
by: Antoniou, Giorgos, et al.
Published: (2025)
Lost-in-the-Middle in Long-Text Generation: Synthetic Dataset, Evaluation Framework, and Mitigation
by: Zhang, Junhao, et al.
Published: (2025)
by: Zhang, Junhao, et al.
Published: (2025)
Conditioning LLMs to Generate Code-Switched Text
by: Heredia, Maite, et al.
Published: (2025)
by: Heredia, Maite, et al.
Published: (2025)
Linguistics Theory Meets LLM: Code-Switched Text Generation via Equivalence Constrained Large Language Models
by: Kuwanto, Garry, et al.
Published: (2024)
by: Kuwanto, Garry, et al.
Published: (2024)
Minimal Pair-Based Evaluation of Code-Switching
by: Sterner, Igor, et al.
Published: (2025)
by: Sterner, Igor, et al.
Published: (2025)
Explaining Predictions by Characteristic Rules
by: Alkhatib, Amr, et al.
Published: (2024)
by: Alkhatib, Amr, et al.
Published: (2024)
LLM Alignment for the Arabs: A Homogenous Culture or Diverse Ones?
by: Keleg, Amr
Published: (2025)
by: Keleg, Amr
Published: (2025)
Evaluating and Achieving Controllable Code Completion in Code LLM
by: Zhang, Jiajun, et al.
Published: (2026)
by: Zhang, Jiajun, et al.
Published: (2026)
OLA: Output Language Alignment in Code-Switched LLM Interactions
by: Oh, Juhyun, et al.
Published: (2026)
by: Oh, Juhyun, et al.
Published: (2026)
AsyncSwitch: Asynchronous Text-Speech Adaptation for Code-Switched ASR
by: Nguyen, Tuan, et al.
Published: (2025)
by: Nguyen, Tuan, et al.
Published: (2025)
Detecting Propaganda Techniques in Code-Switched Social Media Text
by: Salman, Muhammad Umar, et al.
Published: (2023)
by: Salman, Muhammad Umar, et al.
Published: (2023)
Incubating Text Classifiers Following User Instruction with Nothing but LLM
by: Peng, Letian, et al.
Published: (2024)
by: Peng, Letian, et al.
Published: (2024)
PsycoLLM: Enhancing LLM for Psychological Understanding and Evaluation
by: Hu, Jinpeng, et al.
Published: (2024)
by: Hu, Jinpeng, et al.
Published: (2024)
Cell2Text: Multimodal LLM for Generating Single-Cell Descriptions from RNA-Seq Data
by: Kharouiche, Oussama, et al.
Published: (2025)
by: Kharouiche, Oussama, et al.
Published: (2025)
Persona Switch: Mixing Distinct Perspectives in Decoding Time
by: Kim, Junseok, et al.
Published: (2026)
by: Kim, Junseok, et al.
Published: (2026)
Similar Items
-
Fast-Decoding Diffusion Language Models via Progress-Aware Confidence Schedules
by: Mohamed, Amr, et al.
Published: (2025) -
LLM as a Broken Telephone: Iterative Generation Distorts Information
by: Mohamed, Amr, et al.
Published: (2025) -
Beyond Random Sampling: Efficient Language Model Pretraining via Curriculum Learning
by: Zhang, Yang, et al.
Published: (2025) -
The Curious Decline of Linguistic Diversity: Training Language Models on Synthetic Text
by: Guo, Yanzhu, et al.
Published: (2023) -
Markovian Generation Chains in Large Language Models
by: Geng, Mingmeng, et al.
Published: (2026)