Disentangling Codemixing in Chats: The NUS ABC Codemixed Corpus

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Churina, Svetlana, Gupta, Akshat, Mujtahid, Insyirah, Jaidka, Kokil
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913894996377600
author Churina, Svetlana
Gupta, Akshat
Mujtahid, Insyirah
Jaidka, Kokil
author_facet Churina, Svetlana
Gupta, Akshat
Mujtahid, Insyirah
Jaidka, Kokil
contents Code-mixing involves the seamless integration of linguistic elements from multiple languages within a single discourse, reflecting natural multilingual communication patterns. Despite its prominence in informal interactions such as social media, chat messages and instant-messaging exchanges, there has been a lack of publicly available corpora that are author-labeled and suitable for modeling human conversations and relationships. This study introduces the first labeled and general-purpose corpus for understanding code-mixing in context while maintaining rigorous privacy and ethical standards. Our live project will continuously gather, verify, and integrate code-mixed messages into a structured dataset released in JSON format, accompanied by detailed metadata and linguistic statistics. To date, it includes over 355,641 messages spanning various code-mixing patterns, with a primary focus on English, Mandarin, and other languages. We expect the Codemix Corpus to serve as a foundational dataset for research in computational linguistics, sociolinguistics, and NLP applications.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00332
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Disentangling Codemixing in Chats: The NUS ABC Codemixed Corpus
Churina, Svetlana
Gupta, Akshat
Mujtahid, Insyirah
Jaidka, Kokil
Computation and Language
Social and Information Networks
Code-mixing involves the seamless integration of linguistic elements from multiple languages within a single discourse, reflecting natural multilingual communication patterns. Despite its prominence in informal interactions such as social media, chat messages and instant-messaging exchanges, there has been a lack of publicly available corpora that are author-labeled and suitable for modeling human conversations and relationships. This study introduces the first labeled and general-purpose corpus for understanding code-mixing in context while maintaining rigorous privacy and ethical standards. Our live project will continuously gather, verify, and integrate code-mixed messages into a structured dataset released in JSON format, accompanied by detailed metadata and linguistic statistics. To date, it includes over 355,641 messages spanning various code-mixing patterns, with a primary focus on English, Mandarin, and other languages. We expect the Codemix Corpus to serve as a foundational dataset for research in computational linguistics, sociolinguistics, and NLP applications.
title Disentangling Codemixing in Chats: The NUS ABC Codemixed Corpus
topic Computation and Language
Social and Information Networks
url https://arxiv.org/abs/2506.00332