PingPong: A Natural Benchmark for Multi-Turn Code-Switching Dialogues

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Farhansyah, Mohammad Rifqi, Zhafran, Hanif Muhammad, Adilazuarda, Farid, Muhammad, Shamsuddeen Hassan, Mukhtar, Maryam Ibrahim, Ousidhoum, Nedjma, Winata, Genta Indra, Purwarianti, Ayu, Aji, Alham Fikri
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911396090871808
author Farhansyah, Mohammad Rifqi
Zhafran, Hanif Muhammad
Adilazuarda, Farid
Muhammad, Shamsuddeen Hassan
Mukhtar, Maryam Ibrahim
Ousidhoum, Nedjma
Winata, Genta Indra
Purwarianti, Ayu
Aji, Alham Fikri
author_facet Farhansyah, Mohammad Rifqi
Zhafran, Hanif Muhammad
Adilazuarda, Farid
Muhammad, Shamsuddeen Hassan
Mukhtar, Maryam Ibrahim
Ousidhoum, Nedjma
Winata, Genta Indra
Purwarianti, Ayu
Aji, Alham Fikri
contents Code-switching is a widespread practice among the world's multilingual majority, yet few benchmarks accurately reflect its complexity in everyday communication. We present PingPong, a benchmark for natural multi-party code-switching dialogues covering five language-combination variations, some of which are trilingual. Our dataset consists of human-authored conversations among 2 to 4 participants covering authentic, multi-threaded structures where replies frequently reference much earlier points in the dialogue. We demonstrate that our data is significantly more natural and structurally diverse than machine-generated alternatives, offering greater variation in message length, speaker dominance, and reply distance. Based on these dialogues, we define three downstream tasks: Question Answering, Dialogue Summarization, and Topic Classification. Evaluations of several state-of-the-art language models on PingPong reveal that performance remains limited on code-switched inputs, underscoring the urgent need for more robust NLP systems capable of addressing the intricacies of real-world multilingual discourse.
format Preprint
id arxiv_https___arxiv_org_abs_2601_17277
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PingPong: A Natural Benchmark for Multi-Turn Code-Switching Dialogues
Farhansyah, Mohammad Rifqi
Zhafran, Hanif Muhammad
Adilazuarda, Farid
Muhammad, Shamsuddeen Hassan
Mukhtar, Maryam Ibrahim
Ousidhoum, Nedjma
Winata, Genta Indra
Purwarianti, Ayu
Aji, Alham Fikri
Computation and Language
Code-switching is a widespread practice among the world's multilingual majority, yet few benchmarks accurately reflect its complexity in everyday communication. We present PingPong, a benchmark for natural multi-party code-switching dialogues covering five language-combination variations, some of which are trilingual. Our dataset consists of human-authored conversations among 2 to 4 participants covering authentic, multi-threaded structures where replies frequently reference much earlier points in the dialogue. We demonstrate that our data is significantly more natural and structurally diverse than machine-generated alternatives, offering greater variation in message length, speaker dominance, and reply distance. Based on these dialogues, we define three downstream tasks: Question Answering, Dialogue Summarization, and Topic Classification. Evaluations of several state-of-the-art language models on PingPong reveal that performance remains limited on code-switched inputs, underscoring the urgent need for more robust NLP systems capable of addressing the intricacies of real-world multilingual discourse.
title PingPong: A Natural Benchmark for Multi-Turn Code-Switching Dialogues
topic Computation and Language
url https://arxiv.org/abs/2601.17277