Concept-Based Interpretability for Toxicity Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Garg, Samarth, Singh, Divya, Varshney, Deeksha, Mamta |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Indic-TunedLens: Interpreting Multilingual Models in Indian Languages
by: Panchal, Mihir, et al.
Published: (2026)
by: Panchal, Mihir, et al.
Published: (2026)
Towards Robust ESG Analysis Against Greenwashing Risks: Aspect-Action Analysis with Cross-Category Generalization
by: Ong, Keane, et al.
Published: (2025)
by: Ong, Keane, et al.
Published: (2025)
Protein Secondary Structure Prediction Using 3D Graphs and Relation-Aware Message Passing Transformers
by: Varshney, Disha, et al.
Published: (2025)
by: Varshney, Disha, et al.
Published: (2025)
Deriving Strategic Market Insights with Large Language Models: A Benchmark for Forward Counterfactual Generation
by: Ong, Keane, et al.
Published: (2025)
by: Ong, Keane, et al.
Published: (2025)
Cross-Layer Discrete Concept Discovery for Interpreting Language Models
by: Garg, Ankur, et al.
Published: (2025)
by: Garg, Ankur, et al.
Published: (2025)
Concept Based Continuous Prompts for Interpretable Text Classification
by: Chen, Qian, et al.
Published: (2024)
by: Chen, Qian, et al.
Published: (2024)
Memoria: A Scalable Agentic Memory Framework for Personalized Conversational AI
by: Sarin, Samarth, et al.
Published: (2025)
by: Sarin, Samarth, et al.
Published: (2025)
Beyond Toxic: Toxicity Detection Datasets are Not Enough for Brand Safety
by: Korotkova, Elizaveta, et al.
Published: (2023)
by: Korotkova, Elizaveta, et al.
Published: (2023)
Cite Before You Speak: Enhancing Context-Response Grounding in E-commerce Conversational LLM-Agents
by: Zeng, Jingying, et al.
Published: (2025)
by: Zeng, Jingying, et al.
Published: (2025)
Toxicity Begets Toxicity: Unraveling Conversational Chains in Political Podcasts
by: Rizwan, Naquee, et al.
Published: (2025)
by: Rizwan, Naquee, et al.
Published: (2025)
On Theoretical Interpretations of Concept-Based In-Context Learning
by: Tang, Huaze, et al.
Published: (2025)
by: Tang, Huaze, et al.
Published: (2025)
ClimaEmpact: Domain-Aligned Small Language Models and Datasets for Extreme Weather Analytics
by: Varshney, Deeksha, et al.
Published: (2025)
by: Varshney, Deeksha, et al.
Published: (2025)
Enhancing LLM-based Hatred and Toxicity Detection with Meta-Toxic Knowledge Graph
by: Zhao, Yibo, et al.
Published: (2024)
by: Zhao, Yibo, et al.
Published: (2024)
UNITYAI-GUARD: Pioneering Toxicity Detection Across Low-Resource Indian Languages
by: Beniwal, Himanshu, et al.
Published: (2025)
by: Beniwal, Himanshu, et al.
Published: (2025)
Examples as the Prompt: A Scalable Approach for Efficient LLM Adaptation in E-Commerce
by: Zeng, Jingying, et al.
Published: (2025)
by: Zeng, Jingying, et al.
Published: (2025)
Focus on Your Question! Interpreting and Mitigating Toxic CoT Problems in Commonsense Reasoning
by: Li, Jiachun, et al.
Published: (2024)
by: Li, Jiachun, et al.
Published: (2024)
Linearly-Interpretable Concept Embedding Models for Text Analysis
by: De Santis, Francesco, et al.
Published: (2024)
by: De Santis, Francesco, et al.
Published: (2024)
Reliability Analysis of Psychological Concept Extraction and Classification in User-penned Text
by: Garg, Muskan, et al.
Published: (2024)
by: Garg, Muskan, et al.
Published: (2024)
Concept Attractors in LLMs and their Applications
by: Chytas, Sotirios Panagiotis, et al.
Published: (2025)
by: Chytas, Sotirios Panagiotis, et al.
Published: (2025)
ConceptGuard: Neuro-Symbolic Safety Guardrails via Sparse Interpretable Jailbreak Concepts
by: Aswal, Darpan, et al.
Published: (2025)
by: Aswal, Darpan, et al.
Published: (2025)
Thinking About Thinking: Evaluating Reasoning in Post-Trained Language Models
by: Singla, Pratham, et al.
Published: (2025)
by: Singla, Pratham, et al.
Published: (2025)
From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
by: Mueller, Aaron, et al.
Published: (2025)
by: Mueller, Aaron, et al.
Published: (2025)
Palisade -- Prompt Injection Detection Framework
by: Kokkula, Sahasra, et al.
Published: (2024)
by: Kokkula, Sahasra, et al.
Published: (2024)
GloSS over Toxicity: Understanding and Mitigating Toxicity in LLMs via Global Toxic Subspace
by: Duan, Zenghao, et al.
Published: (2025)
by: Duan, Zenghao, et al.
Published: (2025)
Transformer-based Causal Language Models Perform Clustering
by: Wu, Xinbo, et al.
Published: (2024)
by: Wu, Xinbo, et al.
Published: (2024)
Use Sparse Autoencoders to Discover Unknown Concepts, Not to Act on Known Concepts
by: Peng, Kenny, et al.
Published: (2025)
by: Peng, Kenny, et al.
Published: (2025)
ContiGuard: A Framework for Continual Toxicity Detection Against Evolving Evasive Perturbations
by: Kang, Hankun, et al.
Published: (2026)
by: Kang, Hankun, et al.
Published: (2026)
Toxicity Detection is NOT all you Need: Measuring the Gaps to Supporting Volunteer Content Moderators
by: Cao, Yang Trista, et al.
Published: (2023)
by: Cao, Yang Trista, et al.
Published: (2023)
Breaking the Cloak! Unveiling Chinese Cloaked Toxicity with Homophone Graph and Toxic Lexicon
by: Ma, Xuchen, et al.
Published: (2025)
by: Ma, Xuchen, et al.
Published: (2025)
Variational Language Concepts for Interpreting Foundation Language Models
by: Wang, Hengyi, et al.
Published: (2024)
by: Wang, Hengyi, et al.
Published: (2024)
Can LLMs Recognize Toxicity? A Structured Investigation Framework and Toxicity Metric
by: Koh, Hyukhun, et al.
Published: (2024)
by: Koh, Hyukhun, et al.
Published: (2024)
Exploring Multimodal Challenges in Toxic Chinese Detection: Taxonomy, Benchmark, and Findings
by: Yang, Shujian, et al.
Published: (2025)
by: Yang, Shujian, et al.
Published: (2025)
Revisiting Rule-Based Stuttering Detection: A Comprehensive Analysis of Interpretable Models for Clinical Applications
by: Zhang, Eric
Published: (2025)
by: Zhang, Eric
Published: (2025)
xList-Hate: A Checklist-Based Framework for Interpretable and Generalizable Hate Speech Detection
by: Girón, Adrián, et al.
Published: (2026)
by: Girón, Adrián, et al.
Published: (2026)
MUTEX: Leveraging Multilingual Transformers and Conditional Random Fields for Enhanced Urdu Toxic Span Detection
by: Arshad, Inayat, et al.
Published: (2026)
by: Arshad, Inayat, et al.
Published: (2026)
Crafting Narrative Closures: Zero-Shot Learning with SSM Mamba for Short Story Ending Generation
by: Sharma, Divyam, et al.
Published: (2024)
by: Sharma, Divyam, et al.
Published: (2024)
Don't Believe Everything You Read: Enhancing Summarization Interpretability through Automatic Identification of Hallucinations in Large Language Models
by: Vakharia, Priyesh, et al.
Published: (2023)
by: Vakharia, Priyesh, et al.
Published: (2023)
Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models
by: Beniwal, Himanshu, et al.
Published: (2026)
by: Beniwal, Himanshu, et al.
Published: (2026)
Domain Knowledge-Enhanced LLMs for Fraud and Concept Drift Detection
by: Şenol, Ali, et al.
Published: (2025)
by: Şenol, Ali, et al.
Published: (2025)
Agentic Driving Coach: Robustness and Determinism of Agentic AI-Powered Human-in-the-Loop Cyber-Physical Systems
by: Prahlad, Deeksha, et al.
Published: (2026)
by: Prahlad, Deeksha, et al.
Published: (2026)
Similar Items
-
Indic-TunedLens: Interpreting Multilingual Models in Indian Languages
by: Panchal, Mihir, et al.
Published: (2026) -
Towards Robust ESG Analysis Against Greenwashing Risks: Aspect-Action Analysis with Cross-Category Generalization
by: Ong, Keane, et al.
Published: (2025) -
Protein Secondary Structure Prediction Using 3D Graphs and Relation-Aware Message Passing Transformers
by: Varshney, Disha, et al.
Published: (2025) -
Deriving Strategic Market Insights with Large Language Models: A Benchmark for Forward Counterfactual Generation
by: Ong, Keane, et al.
Published: (2025) -
Cross-Layer Discrete Concept Discovery for Interpreting Language Models
by: Garg, Ankur, et al.
Published: (2025)