MixSarc: A Bangla-English Code-Mixed Corpus for Implicit Meaning Identification

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Alam, Kazi Samin Yasar, Chowdhury, Md Tanbir, Ahmed, Tamim, Abrar, Ajwad, Haque, Md Rafid
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914350222016512
author Alam, Kazi Samin Yasar
Chowdhury, Md Tanbir
Ahmed, Tamim
Abrar, Ajwad
Haque, Md Rafid
author_facet Alam, Kazi Samin Yasar
Chowdhury, Md Tanbir
Ahmed, Tamim
Abrar, Ajwad
Haque, Md Rafid
contents Bangla-English code-mixing is widespread across South Asian social media, yet resources for implicit meaning identification in this setting remain scarce. Existing sentiment and sarcasm models largely focus on monolingual English or high-resource languages and struggle with transliteration variation, cultural references, and intra-sentential language switching. To address this gap, we introduce MixSarc, the first publicly available Bangla-English code-mixed corpus for implicit meaning identification. The dataset contains 9,087 manually annotated sentences labeled for humor, sarcasm, offensiveness, and vulgarity. We construct the corpus through targeted social media collection, systematic filtering, and multi-annotator validation. We benchmark transformer-based models and evaluate zero-shot large language models under structured prompting. Results show strong performance on humor detection but substantial degradation on sarcasm, offense, and vulgarity due to class imbalance and pragmatic complexity. Zero-shot models achieve competitive micro-F1 scores but low exact match accuracy. Further analysis reveals that over 42\% of negative sentiment instances in an external dataset exhibit sarcastic characteristics. MixSarc provides a foundational resource for culturally aware NLP and supports more reliable multi-label modeling in code-mixed environments.
format Preprint
id arxiv_https___arxiv_org_abs_2602_21608
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MixSarc: A Bangla-English Code-Mixed Corpus for Implicit Meaning Identification
Alam, Kazi Samin Yasar
Chowdhury, Md Tanbir
Ahmed, Tamim
Abrar, Ajwad
Haque, Md Rafid
Computation and Language
Bangla-English code-mixing is widespread across South Asian social media, yet resources for implicit meaning identification in this setting remain scarce. Existing sentiment and sarcasm models largely focus on monolingual English or high-resource languages and struggle with transliteration variation, cultural references, and intra-sentential language switching. To address this gap, we introduce MixSarc, the first publicly available Bangla-English code-mixed corpus for implicit meaning identification. The dataset contains 9,087 manually annotated sentences labeled for humor, sarcasm, offensiveness, and vulgarity. We construct the corpus through targeted social media collection, systematic filtering, and multi-annotator validation. We benchmark transformer-based models and evaluate zero-shot large language models under structured prompting. Results show strong performance on humor detection but substantial degradation on sarcasm, offense, and vulgarity due to class imbalance and pragmatic complexity. Zero-shot models achieve competitive micro-F1 scores but low exact match accuracy. Further analysis reveals that over 42\% of negative sentiment instances in an external dataset exhibit sarcastic characteristics. MixSarc provides a foundational resource for culturally aware NLP and supports more reliable multi-label modeling in code-mixed environments.
title MixSarc: A Bangla-English Code-Mixed Corpus for Implicit Meaning Identification
topic Computation and Language
url https://arxiv.org/abs/2602.21608