Toxicity Begets Toxicity: Unraveling Conversational Chains in Political Podcasts
Fuente:
arXiv
Saved in:
| Main Authors: | Rizwan, Naquee, Deb, Nayandeep, Roy, Sarthak, Solanki, Vishwajeet Singh, Garimella, Kiran, Mukherjee, Animesh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
STEMTOX: From Social Tags to Fine-Grained Toxic Meme Detection via Entropy-Guided Multi-Task Learning
by: Swain, Subhankar, et al.
Published: (2025)
by: Swain, Subhankar, et al.
Published: (2025)
FBHM: Functional Benchmarking and Steering of VLMs for Hateful Meme Detection
by: Bhaskar, Paramananda, et al.
Published: (2026)
by: Bhaskar, Paramananda, et al.
Published: (2026)
Concept-Based Interpretability for Toxicity Detection
by: Garg, Samarth, et al.
Published: (2025)
by: Garg, Samarth, et al.
Published: (2025)
GloSS over Toxicity: Understanding and Mitigating Toxicity in LLMs via Global Toxic Subspace
by: Duan, Zenghao, et al.
Published: (2025)
by: Duan, Zenghao, et al.
Published: (2025)
Beyond Toxic: Toxicity Detection Datasets are Not Enough for Brand Safety
by: Korotkova, Elizaveta, et al.
Published: (2023)
by: Korotkova, Elizaveta, et al.
Published: (2023)
Breaking the Cloak! Unveiling Chinese Cloaked Toxicity with Homophone Graph and Toxic Lexicon
by: Ma, Xuchen, et al.
Published: (2025)
by: Ma, Xuchen, et al.
Published: (2025)
Can LLMs Recognize Toxicity? A Structured Investigation Framework and Toxicity Metric
by: Koh, Hyukhun, et al.
Published: (2024)
by: Koh, Hyukhun, et al.
Published: (2024)
Enhancing LLM-based Hatred and Toxicity Detection with Meta-Toxic Knowledge Graph
by: Zhao, Yibo, et al.
Published: (2024)
by: Zhao, Yibo, et al.
Published: (2024)
Pragmatic Inference Chain (PIC) Improving LLMs' Reasoning of Authentic Implicit Toxic Language
by: Chen, Xi, et al.
Published: (2025)
by: Chen, Xi, et al.
Published: (2025)
Multilingual and Explainable Text Detoxification with Parallel Corpora
by: Dementieva, Daryna, et al.
Published: (2024)
by: Dementieva, Daryna, et al.
Published: (2024)
IndoToxic2024: A Demographically-Enriched Dataset of Hate Speech and Toxicity Types for Indonesian Language
by: Susanto, Lucky, et al.
Published: (2024)
by: Susanto, Lucky, et al.
Published: (2024)
Toxicity-Aware Few-Shot Prompting for Low-Resource Singlish Translation
by: Ge, Ziyu, et al.
Published: (2025)
by: Ge, Ziyu, et al.
Published: (2025)
Characterising Toxicity in Generative Large Language Models
by: Zhang, Zhiyao, et al.
Published: (2026)
by: Zhang, Zhiyao, et al.
Published: (2026)
Realistic Evaluation of Toxicity in Large Language Models
by: Luong, Tinh Son, et al.
Published: (2024)
by: Luong, Tinh Son, et al.
Published: (2024)
Characterization of Political Polarized Users Attacked by Language Toxicity on Twitter
by: Xu, Wentao
Published: (2024)
by: Xu, Wentao
Published: (2024)
UNITYAI-GUARD: Pioneering Toxicity Detection Across Low-Resource Indian Languages
by: Beniwal, Himanshu, et al.
Published: (2025)
by: Beniwal, Himanshu, et al.
Published: (2025)
Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models
by: Beniwal, Himanshu, et al.
Published: (2026)
by: Beniwal, Himanshu, et al.
Published: (2026)
How Toxic Can You Get? Search-based Toxicity Testing for Large Language Models
by: Corbo, Simone, et al.
Published: (2025)
by: Corbo, Simone, et al.
Published: (2025)
Whispering Experts: Neural Interventions for Toxicity Mitigation in Language Models
by: Suau, Xavier, et al.
Published: (2024)
by: Suau, Xavier, et al.
Published: (2024)
On the Relationship between Truth and Political Bias in Language Models
by: Fulay, Suyash, et al.
Published: (2024)
by: Fulay, Suyash, et al.
Published: (2024)
Something Just Like TRuST : Toxicity Recognition of Span and Target
by: Atil, Berk, et al.
Published: (2025)
by: Atil, Berk, et al.
Published: (2025)
MDIT-Bench: Evaluating the Dual-Implicit Toxicity in Large Multimodal Models
by: Jin, Bohan, et al.
Published: (2025)
by: Jin, Bohan, et al.
Published: (2025)
From One to Many: Expanding the Scope of Toxicity Mitigation in Language Models
by: Pozzobon, Luiza, et al.
Published: (2024)
by: Pozzobon, Luiza, et al.
Published: (2024)
Bridging Context Gaps: Enhancing Comprehension in Long-Form Social Conversations Through Contextualized Excerpts
by: Mohanty, Shrestha, et al.
Published: (2024)
by: Mohanty, Shrestha, et al.
Published: (2024)
Rethinking Toxicity Evaluation in Large Language Models: A Multi-Label Perspective
by: Kou, Zhiqiang, et al.
Published: (2025)
by: Kou, Zhiqiang, et al.
Published: (2025)
Induction Head Toxicity Mechanistically Explains Repetition Curse in Large Language Models
by: Wang, Shuxun, et al.
Published: (2025)
by: Wang, Shuxun, et al.
Published: (2025)
A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
by: Lee, Andrew, et al.
Published: (2024)
by: Lee, Andrew, et al.
Published: (2024)
Focus on Your Question! Interpreting and Mitigating Toxic CoT Problems in Commonsense Reasoning
by: Li, Jiachun, et al.
Published: (2024)
by: Li, Jiachun, et al.
Published: (2024)
See, Explain, and Intervene: A Few-Shot Multimodal Agent Framework for Hateful Meme Moderation
by: Rizwan, Naquee, et al.
Published: (2026)
by: Rizwan, Naquee, et al.
Published: (2026)
Exploring Multimodal Challenges in Toxic Chinese Detection: Taxonomy, Benchmark, and Findings
by: Yang, Shujian, et al.
Published: (2025)
by: Yang, Shujian, et al.
Published: (2025)
Polarized Patterns of Language Toxicity and Sentiment of Debunking Posts on Social Media
by: Xu, Wentao, et al.
Published: (2025)
by: Xu, Wentao, et al.
Published: (2025)
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation
by: Zhang, Zhibo, et al.
Published: (2025)
by: Zhang, Zhibo, et al.
Published: (2025)
ContiGuard: A Framework for Continual Toxicity Detection Against Evolving Evasive Perturbations
by: Kang, Hankun, et al.
Published: (2026)
by: Kang, Hankun, et al.
Published: (2026)
Toxicity Detection is NOT all you Need: Measuring the Gaps to Supporting Volunteer Content Moderators
by: Cao, Yang Trista, et al.
Published: (2023)
by: Cao, Yang Trista, et al.
Published: (2023)
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions
by: Xu, Jingxin, et al.
Published: (2025)
by: Xu, Jingxin, et al.
Published: (2025)
Building Resource-Constrained Language Agents: A Korean Case Study on Chemical Toxicity Information
by: Cho, Hojun, et al.
Published: (2025)
by: Cho, Hojun, et al.
Published: (2025)
Toxic Memes: A Survey of Computational Perspectives on the Detection and Explanation of Meme Toxicities
by: Pandiani, Delfina Sol Martinez, et al.
Published: (2024)
by: Pandiani, Delfina Sol Martinez, et al.
Published: (2024)
ToxiLab: How Well Do Open-Source LLMs Generate Synthetic Toxicity Data?
by: Hui, Zheng, et al.
Published: (2024)
by: Hui, Zheng, et al.
Published: (2024)
MUTEX: Leveraging Multilingual Transformers and Conditional Random Fields for Enhanced Urdu Toxic Span Detection
by: Arshad, Inayat, et al.
Published: (2026)
by: Arshad, Inayat, et al.
Published: (2026)
Adopt $\neq$ Adapt: Longitudinal Analyses of LLM Conversations in the Wild
by: Hicke, Rebecca M. M., et al.
Published: (2026)
by: Hicke, Rebecca M. M., et al.
Published: (2026)
Similar Items
-
STEMTOX: From Social Tags to Fine-Grained Toxic Meme Detection via Entropy-Guided Multi-Task Learning
by: Swain, Subhankar, et al.
Published: (2025) -
FBHM: Functional Benchmarking and Steering of VLMs for Hateful Meme Detection
by: Bhaskar, Paramananda, et al.
Published: (2026) -
Concept-Based Interpretability for Toxicity Detection
by: Garg, Samarth, et al.
Published: (2025) -
GloSS over Toxicity: Understanding and Mitigating Toxicity in LLMs via Global Toxic Subspace
by: Duan, Zenghao, et al.
Published: (2025) -
Beyond Toxic: Toxicity Detection Datasets are Not Enough for Brand Safety
by: Korotkova, Elizaveta, et al.
Published: (2023)