CS-Sum: A Benchmark for Code-Switching Dialogue Summarization and the Limits of Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Suresh, Sathya Krishnan, Surana, Tanmay, Hao, Lim Zhi, Chng, Eng Siong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916745119268864
author Suresh, Sathya Krishnan
Surana, Tanmay
Hao, Lim Zhi
Chng, Eng Siong
author_facet Suresh, Sathya Krishnan
Surana, Tanmay
Hao, Lim Zhi
Chng, Eng Siong
contents Code-switching (CS) poses a significant challenge for Large Language Models (LLMs), yet its comprehensibility remains underexplored in LLMs. We introduce CS-Sum, to evaluate the comprehensibility of CS by the LLMs through CS dialogue to English summarization. CS-Sum is the first benchmark for CS dialogue summarization across Mandarin-English (EN-ZH), Tamil-English (EN-TA), and Malay-English (EN-MS), with 900-1300 human-annotated dialogues per language pair. Evaluating ten LLMs, including open and closed-source models, we analyze performance across few-shot, translate-summarize, and fine-tuning (LoRA, QLoRA on synthetic data) approaches. Our findings show that though the scores on automated metrics are high, LLMs make subtle mistakes that alter the complete meaning of the dialogue. To this end, we introduce 3 most common type of errors that LLMs make when handling CS input. Error rates vary across CS pairs and LLMs, with some LLMs showing more frequent errors on certain language pairs, underscoring the need for specialized training on code-switched data.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13559
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CS-Sum: A Benchmark for Code-Switching Dialogue Summarization and the Limits of Large Language Models
Suresh, Sathya Krishnan
Surana, Tanmay
Hao, Lim Zhi
Chng, Eng Siong
Computation and Language
Machine Learning
Code-switching (CS) poses a significant challenge for Large Language Models (LLMs), yet its comprehensibility remains underexplored in LLMs. We introduce CS-Sum, to evaluate the comprehensibility of CS by the LLMs through CS dialogue to English summarization. CS-Sum is the first benchmark for CS dialogue summarization across Mandarin-English (EN-ZH), Tamil-English (EN-TA), and Malay-English (EN-MS), with 900-1300 human-annotated dialogues per language pair. Evaluating ten LLMs, including open and closed-source models, we analyze performance across few-shot, translate-summarize, and fine-tuning (LoRA, QLoRA on synthetic data) approaches. Our findings show that though the scores on automated metrics are high, LLMs make subtle mistakes that alter the complete meaning of the dialogue. To this end, we introduce 3 most common type of errors that LLMs make when handling CS input. Error rates vary across CS pairs and LLMs, with some LLMs showing more frequent errors on certain language pairs, underscoring the need for specialized training on code-switched data.
title CS-Sum: A Benchmark for Code-Switching Dialogue Summarization and the Limits of Large Language Models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2505.13559