ToxiLab: How Well Do Open-Source LLMs Generate Synthetic Toxicity Data?

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hui, Zheng, Guo, Zhaoxiao, Zhao, Hang, Duan, Juanyong, Ai, Lin, Li, Yinheng, Hirschberg, Julia, Huang, Congrui
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917933186285568
author Hui, Zheng
Guo, Zhaoxiao
Zhao, Hang
Duan, Juanyong
Ai, Lin
Li, Yinheng
Hirschberg, Julia
Huang, Congrui
author_facet Hui, Zheng
Guo, Zhaoxiao
Zhao, Hang
Duan, Juanyong
Ai, Lin
Li, Yinheng
Hirschberg, Julia
Huang, Congrui
contents Effective toxic content detection relies heavily on high-quality and diverse data, which serve as the foundation for robust content moderation models. Synthetic data has become a common approach for training models across various NLP tasks. However, its effectiveness remains uncertain for highly subjective tasks like hate speech detection, with previous research yielding mixed results. This study explores the potential of open-source LLMs for harmful data synthesis, utilizing controlled prompting and supervised fine-tuning techniques to enhance data quality and diversity. We systematically evaluated 6 open source LLMs on 5 datasets, assessing their ability to generate diverse, high-quality harmful data while minimizing hallucination and duplication. Our results show that Mistral consistently outperforms other open models, and supervised fine-tuning significantly enhances data reliability and diversity. We further analyze the trade-offs between prompt-based vs. fine-tuned toxic data synthesis, discuss real-world deployment challenges, and highlight ethical considerations. Our findings demonstrate that fine-tuned open source LLMs provide scalable and cost-effective solutions to augment toxic content detection datasets, paving the way for more accessible and transparent content moderation tools.
format Preprint
id arxiv_https___arxiv_org_abs_2411_15175
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ToxiLab: How Well Do Open-Source LLMs Generate Synthetic Toxicity Data?
Hui, Zheng
Guo, Zhaoxiao
Zhao, Hang
Duan, Juanyong
Ai, Lin
Li, Yinheng
Hirschberg, Julia
Huang, Congrui
Computation and Language
Artificial Intelligence
Effective toxic content detection relies heavily on high-quality and diverse data, which serve as the foundation for robust content moderation models. Synthetic data has become a common approach for training models across various NLP tasks. However, its effectiveness remains uncertain for highly subjective tasks like hate speech detection, with previous research yielding mixed results. This study explores the potential of open-source LLMs for harmful data synthesis, utilizing controlled prompting and supervised fine-tuning techniques to enhance data quality and diversity. We systematically evaluated 6 open source LLMs on 5 datasets, assessing their ability to generate diverse, high-quality harmful data while minimizing hallucination and duplication. Our results show that Mistral consistently outperforms other open models, and supervised fine-tuning significantly enhances data reliability and diversity. We further analyze the trade-offs between prompt-based vs. fine-tuned toxic data synthesis, discuss real-world deployment challenges, and highlight ethical considerations. Our findings demonstrate that fine-tuned open source LLMs provide scalable and cost-effective solutions to augment toxic content detection datasets, paving the way for more accessible and transparent content moderation tools.
title ToxiLab: How Well Do Open-Source LLMs Generate Synthetic Toxicity Data?
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2411.15175