Toxicity-Aware Few-Shot Prompting for Low-Resource Singlish Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ge, Ziyu, Chua, Gabriel, Tan, Leanne, Lee, Roy Ka-Wei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911059442401280
author Ge, Ziyu
Chua, Gabriel
Tan, Leanne
Lee, Roy Ka-Wei
author_facet Ge, Ziyu
Chua, Gabriel
Tan, Leanne
Lee, Roy Ka-Wei
contents As online communication increasingly incorporates under-represented languages and colloquial dialects, standard translation systems often fail to preserve local slang, code-mixing, and culturally embedded markers of harmful speech. Translating toxic content between low-resource language pairs poses additional challenges due to scarce parallel data and safety filters that sanitize offensive expressions. In this work, we propose a reproducible, two-stage framework for toxicity-preserving translation, demonstrated on a code-mixed Singlish safety corpus. First, we perform human-verified few-shot prompt engineering: we iteratively curate and rank annotator-selected Singlish-target examples to capture nuanced slang, tone, and toxicity. Second, we optimize model-prompt pairs by benchmarking several large language models using semantic similarity via direct and back-translation. Quantitative human evaluation confirms the effectiveness and efficiency of our pipeline. Beyond improving translation quality, our framework contributes to the safety of multicultural LLMs by supporting culturally sensitive moderation and benchmarking in low-resource contexts. By positioning Singlish as a testbed for inclusive NLP, we underscore the importance of preserving sociolinguistic nuance in real-world applications such as content moderation and regional platform governance.
format Preprint
id arxiv_https___arxiv_org_abs_2507_11966
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Toxicity-Aware Few-Shot Prompting for Low-Resource Singlish Translation
Ge, Ziyu
Chua, Gabriel
Tan, Leanne
Lee, Roy Ka-Wei
Computation and Language
Artificial Intelligence
Computers and Society
As online communication increasingly incorporates under-represented languages and colloquial dialects, standard translation systems often fail to preserve local slang, code-mixing, and culturally embedded markers of harmful speech. Translating toxic content between low-resource language pairs poses additional challenges due to scarce parallel data and safety filters that sanitize offensive expressions. In this work, we propose a reproducible, two-stage framework for toxicity-preserving translation, demonstrated on a code-mixed Singlish safety corpus. First, we perform human-verified few-shot prompt engineering: we iteratively curate and rank annotator-selected Singlish-target examples to capture nuanced slang, tone, and toxicity. Second, we optimize model-prompt pairs by benchmarking several large language models using semantic similarity via direct and back-translation. Quantitative human evaluation confirms the effectiveness and efficiency of our pipeline. Beyond improving translation quality, our framework contributes to the safety of multicultural LLMs by supporting culturally sensitive moderation and benchmarking in low-resource contexts. By positioning Singlish as a testbed for inclusive NLP, we underscore the importance of preserving sociolinguistic nuance in real-world applications such as content moderation and regional platform governance.
title Toxicity-Aware Few-Shot Prompting for Low-Resource Singlish Translation
topic Computation and Language
Artificial Intelligence
Computers and Society
url https://arxiv.org/abs/2507.11966