Dialectal Toxicity Detection: Evaluating LLM-as-a-Judge Consistency Across Language Varieties

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Faisal, Fahim, Rahman, Md Mushfiqur, Anastasopoulos, Antonios
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912122630307840
author Faisal, Fahim
Rahman, Md Mushfiqur
Anastasopoulos, Antonios
author_facet Faisal, Fahim
Rahman, Md Mushfiqur
Anastasopoulos, Antonios
contents There has been little systematic study on how dialectal differences affect toxicity detection by modern LLMs. Furthermore, although using LLMs as evaluators ("LLM-as-a-judge") is a growing research area, their sensitivity to dialectal nuances is still underexplored and requires more focused attention. In this paper, we address these gaps through a comprehensive toxicity evaluation of LLMs across diverse dialects. We create a multi-dialect dataset through synthetic transformations and human-assisted translations, covering 10 language clusters and 60 varieties. We then evaluated three LLMs on their ability to assess toxicity across multilingual, dialectal, and LLM-human consistency. Our findings show that LLMs are sensitive in handling both multilingual and dialectal variations. However, if we have to rank the consistency, the weakest area is LLM-human agreement, followed by dialectal consistency. Code repository: \url{https://github.com/ffaisal93/dialect_toxicity_llm_judge}
format Preprint
id arxiv_https___arxiv_org_abs_2411_10954
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Dialectal Toxicity Detection: Evaluating LLM-as-a-Judge Consistency Across Language Varieties
Faisal, Fahim
Rahman, Md Mushfiqur
Anastasopoulos, Antonios
Computation and Language
There has been little systematic study on how dialectal differences affect toxicity detection by modern LLMs. Furthermore, although using LLMs as evaluators ("LLM-as-a-judge") is a growing research area, their sensitivity to dialectal nuances is still underexplored and requires more focused attention. In this paper, we address these gaps through a comprehensive toxicity evaluation of LLMs across diverse dialects. We create a multi-dialect dataset through synthetic transformations and human-assisted translations, covering 10 language clusters and 60 varieties. We then evaluated three LLMs on their ability to assess toxicity across multilingual, dialectal, and LLM-human consistency. Our findings show that LLMs are sensitive in handling both multilingual and dialectal variations. However, if we have to rank the consistency, the weakest area is LLM-human agreement, followed by dialectal consistency. Code repository: \url{https://github.com/ffaisal93/dialect_toxicity_llm_judge}
title Dialectal Toxicity Detection: Evaluating LLM-as-a-Judge Consistency Across Language Varieties
topic Computation and Language
url https://arxiv.org/abs/2411.10954