Beyond Toxic: Toxicity Detection Datasets are Not Enough for Brand Safety
Fuente:
arXiv
Saved in:
| Main Authors: | Korotkova, Elizaveta, Chung, Isaac |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhancing LLM-based Hatred and Toxicity Detection with Meta-Toxic Knowledge Graph
by: Zhao, Yibo, et al.
Published: (2024)
by: Zhao, Yibo, et al.
Published: (2024)
IndoToxic2024: A Demographically-Enriched Dataset of Hate Speech and Toxicity Types for Indonesian Language
by: Susanto, Lucky, et al.
Published: (2024)
by: Susanto, Lucky, et al.
Published: (2024)
Concept-Based Interpretability for Toxicity Detection
by: Garg, Samarth, et al.
Published: (2025)
by: Garg, Samarth, et al.
Published: (2025)
GloSS over Toxicity: Understanding and Mitigating Toxicity in LLMs via Global Toxic Subspace
by: Duan, Zenghao, et al.
Published: (2025)
by: Duan, Zenghao, et al.
Published: (2025)
Toxicity Begets Toxicity: Unraveling Conversational Chains in Political Podcasts
by: Rizwan, Naquee, et al.
Published: (2025)
by: Rizwan, Naquee, et al.
Published: (2025)
Can LLMs Recognize Toxicity? A Structured Investigation Framework and Toxicity Metric
by: Koh, Hyukhun, et al.
Published: (2024)
by: Koh, Hyukhun, et al.
Published: (2024)
Breaking the Cloak! Unveiling Chinese Cloaked Toxicity with Homophone Graph and Toxic Lexicon
by: Ma, Xuchen, et al.
Published: (2025)
by: Ma, Xuchen, et al.
Published: (2025)
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation
by: Zhang, Zhibo, et al.
Published: (2025)
by: Zhang, Zhibo, et al.
Published: (2025)
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions
by: Xu, Jingxin, et al.
Published: (2025)
by: Xu, Jingxin, et al.
Published: (2025)
UNITYAI-GUARD: Pioneering Toxicity Detection Across Low-Resource Indian Languages
by: Beniwal, Himanshu, et al.
Published: (2025)
by: Beniwal, Himanshu, et al.
Published: (2025)
Exploring Multimodal Challenges in Toxic Chinese Detection: Taxonomy, Benchmark, and Findings
by: Yang, Shujian, et al.
Published: (2025)
by: Yang, Shujian, et al.
Published: (2025)
Characterising Toxicity in Generative Large Language Models
by: Zhang, Zhiyao, et al.
Published: (2026)
by: Zhang, Zhiyao, et al.
Published: (2026)
Realistic Evaluation of Toxicity in Large Language Models
by: Luong, Tinh Son, et al.
Published: (2024)
by: Luong, Tinh Son, et al.
Published: (2024)
Toxicity Detection is NOT all you Need: Measuring the Gaps to Supporting Volunteer Content Moderators
by: Cao, Yang Trista, et al.
Published: (2023)
by: Cao, Yang Trista, et al.
Published: (2023)
ContiGuard: A Framework for Continual Toxicity Detection Against Evolving Evasive Perturbations
by: Kang, Hankun, et al.
Published: (2026)
by: Kang, Hankun, et al.
Published: (2026)
Towards Detecting Contextual Real-Time Toxicity for In-Game Chat
by: Yang, Zachary, et al.
Published: (2023)
by: Yang, Zachary, et al.
Published: (2023)
Toxic Memes: A Survey of Computational Perspectives on the Detection and Explanation of Meme Toxicities
by: Pandiani, Delfina Sol Martinez, et al.
Published: (2024)
by: Pandiani, Delfina Sol Martinez, et al.
Published: (2024)
How Toxic Can You Get? Search-based Toxicity Testing for Large Language Models
by: Corbo, Simone, et al.
Published: (2025)
by: Corbo, Simone, et al.
Published: (2025)
Toxicity Detection towards Adaptability to Changing Perturbations
by: Kang, Hankun, et al.
Published: (2024)
by: Kang, Hankun, et al.
Published: (2024)
MUTEX: Leveraging Multilingual Transformers and Conditional Random Fields for Enhanced Urdu Toxic Span Detection
by: Arshad, Inayat, et al.
Published: (2026)
by: Arshad, Inayat, et al.
Published: (2026)
Toxicity Detection Should Measure Contextual Harm, Not Text-Intrinsic Badness
by: Berezin, Sergei, et al.
Published: (2025)
by: Berezin, Sergei, et al.
Published: (2025)
Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content
by: Stepanov, Ihor, et al.
Published: (2026)
by: Stepanov, Ihor, et al.
Published: (2026)
Whispering Experts: Neural Interventions for Toxicity Mitigation in Language Models
by: Suau, Xavier, et al.
Published: (2024)
by: Suau, Xavier, et al.
Published: (2024)
Efficient Detection of Toxic Prompts in Large Language Models
by: Liu, Yi, et al.
Published: (2024)
by: Liu, Yi, et al.
Published: (2024)
Obfuscation Rules for Detecting and Detoxifying Korean Toxicity
by: Lee, Yejin, et al.
Published: (2025)
by: Lee, Yejin, et al.
Published: (2025)
Characterizing Large Language Model Geometry Helps Solve Toxicity Detection and Generation
by: Balestriero, Randall, et al.
Published: (2023)
by: Balestriero, Randall, et al.
Published: (2023)
Take its Essence, Discard its Dross! Debiasing for Toxic Language Detection via Counterfactual Causal Effect
by: Lu, Junyu, et al.
Published: (2024)
by: Lu, Junyu, et al.
Published: (2024)
Something Just Like TRuST : Toxicity Recognition of Span and Target
by: Atil, Berk, et al.
Published: (2025)
by: Atil, Berk, et al.
Published: (2025)
MDIT-Bench: Evaluating the Dual-Implicit Toxicity in Large Multimodal Models
by: Jin, Bohan, et al.
Published: (2025)
by: Jin, Bohan, et al.
Published: (2025)
From One to Many: Expanding the Scope of Toxicity Mitigation in Language Models
by: Pozzobon, Luiza, et al.
Published: (2024)
by: Pozzobon, Luiza, et al.
Published: (2024)
SMARTER: A Data-efficient Framework to Improve Toxicity Detection with Explanation via Self-augmenting Large Language Models
by: Nghiem, Huy, et al.
Published: (2025)
by: Nghiem, Huy, et al.
Published: (2025)
Evading Toxicity Detection with ASCII-art: A Benchmark of Spatial Attacks on Moderation Systems
by: Berezin, Sergey, et al.
Published: (2024)
by: Berezin, Sergey, et al.
Published: (2024)
Rethinking Toxicity Evaluation in Large Language Models: A Multi-Label Perspective
by: Kou, Zhiqiang, et al.
Published: (2025)
by: Kou, Zhiqiang, et al.
Published: (2025)
A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
by: Lee, Andrew, et al.
Published: (2024)
by: Lee, Andrew, et al.
Published: (2024)
Focus on Your Question! Interpreting and Mitigating Toxic CoT Problems in Commonsense Reasoning
by: Li, Jiachun, et al.
Published: (2024)
by: Li, Jiachun, et al.
Published: (2024)
Induction Head Toxicity Mechanistically Explains Repetition Curse in Large Language Models
by: Wang, Shuxun, et al.
Published: (2025)
by: Wang, Shuxun, et al.
Published: (2025)
Polarized Patterns of Language Toxicity and Sentiment of Debunking Posts on Social Media
by: Xu, Wentao, et al.
Published: (2025)
by: Xu, Wentao, et al.
Published: (2025)
Toxicity-Aware Few-Shot Prompting for Low-Resource Singlish Translation
by: Ge, Ziyu, et al.
Published: (2025)
by: Ge, Ziyu, et al.
Published: (2025)
Pragmatic Inference Chain (PIC) Improving LLMs' Reasoning of Authentic Implicit Toxic Language
by: Chen, Xi, et al.
Published: (2025)
by: Chen, Xi, et al.
Published: (2025)
ToxVidLM: A Multimodal Framework for Toxicity Detection in Code-Mixed Videos
by: Maity, Krishanu, et al.
Published: (2024)
by: Maity, Krishanu, et al.
Published: (2024)
Similar Items
-
Enhancing LLM-based Hatred and Toxicity Detection with Meta-Toxic Knowledge Graph
by: Zhao, Yibo, et al.
Published: (2024) -
IndoToxic2024: A Demographically-Enriched Dataset of Hate Speech and Toxicity Types for Indonesian Language
by: Susanto, Lucky, et al.
Published: (2024) -
Concept-Based Interpretability for Toxicity Detection
by: Garg, Samarth, et al.
Published: (2025) -
GloSS over Toxicity: Understanding and Mitigating Toxicity in LLMs via Global Toxic Subspace
by: Duan, Zenghao, et al.
Published: (2025) -
Toxicity Begets Toxicity: Unraveling Conversational Chains in Political Podcasts
by: Rizwan, Naquee, et al.
Published: (2025)