Argument-Based Consistency in Toxicity Explanations of LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Mothilal, Ramaravind Kommiya, Roy, Joanna, Ahmed, Syed Ishtiaque, Guha, Shion |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards a Non-Ideal Methodological Framework for Responsible ML
by: Mothilal, Ramaravind Kommiya, et al.
Published: (2024)
by: Mothilal, Ramaravind Kommiya, et al.
Published: (2024)
Reasoning About Reasoning: Towards Informed and Reflective Use of LLM Reasoning in HCI
by: Mothilal, Ramaravind Kommiya, et al.
Published: (2025)
by: Mothilal, Ramaravind Kommiya, et al.
Published: (2025)
Talking About the Assumption in the Room
by: Mothilal, Ramaravind Kommiya, et al.
Published: (2025)
by: Mothilal, Ramaravind Kommiya, et al.
Published: (2025)
BTPD: A Multilingual Hand-curated Dataset of Bengali Transnational Political Discourse Across Online Communities
by: Das, Dipto, et al.
Published: (2025)
by: Das, Dipto, et al.
Published: (2025)
Bureaucratic Silences: What the Canadian AI Register Reveals, Omits, and Obscures
by: Das, Dipto, et al.
Published: (2026)
by: Das, Dipto, et al.
Published: (2026)
How do the Global South Diasporas Mobilize for Transnational Political Change?
by: Das, Dipto, et al.
Published: (2026)
by: Das, Dipto, et al.
Published: (2026)
Attention Consistency for LLMs Explanation
by: Lan, Tian, et al.
Published: (2025)
by: Lan, Tian, et al.
Published: (2025)
How do datasets, developers, and models affect biases in a low-resourced language?: The Case of the Bengali Language
by: Das, Dipto, et al.
Published: (2025)
by: Das, Dipto, et al.
Published: (2025)
The "Colonial Impulse" of Natural Language Processing: An Audit of Bengali Sentiment Analysis Tools and Their Identity-based Biases
by: Das, Dipto, et al.
Published: (2024)
by: Das, Dipto, et al.
Published: (2024)
Aligning What LLMs Do and Say: Towards Self-Consistent Explanations
by: Admoni, Sahar, et al.
Published: (2025)
by: Admoni, Sahar, et al.
Published: (2025)
Main Predicate and Their Arguments as Explanation Signals For Intent Classification
by: Pimparkhede, Sameer, et al.
Published: (2025)
by: Pimparkhede, Sameer, et al.
Published: (2025)
FineLIP: Extending CLIP's Reach via Fine-Grained Alignment with Longer Text Inputs
by: Asokan, Mothilal, et al.
Published: (2025)
by: Asokan, Mothilal, et al.
Published: (2025)
Are You the A-hole? A Fair, Multi-Perspective Ethical Reasoning Framework
by: Munir, Sheza, et al.
Published: (2026)
by: Munir, Sheza, et al.
Published: (2026)
Understanding How CodeLLMs (Mis)Predict Types with Activation Steering
by: Lucchetti, Francesca, et al.
Published: (2024)
by: Lucchetti, Francesca, et al.
Published: (2024)
Can LLMs Extract Frame-Semantic Arguments?
by: Devasier, Jacob, et al.
Published: (2025)
by: Devasier, Jacob, et al.
Published: (2025)
Argument Reconstruction as Supervision for Critical Thinking in LLMs
by: Ryu, Hyun, et al.
Published: (2026)
by: Ryu, Hyun, et al.
Published: (2026)
LLMs for Argument Mining: Detection, Extraction, and Relationship Classification of pre-defined Arguments in Online Comments
by: Guida, Matteo, et al.
Published: (2025)
by: Guida, Matteo, et al.
Published: (2025)
LLM-Based Multi-Task Bangla Hate Speech Detection: Type, Severity, and Target
by: Hasan, Md Arid, et al.
Published: (2025)
by: Hasan, Md Arid, et al.
Published: (2025)
Towards Consistent Natural-Language Explanations via Explanation-Consistency Finetuning
by: Chen, Yanda, et al.
Published: (2024)
by: Chen, Yanda, et al.
Published: (2024)
Leveraging Small LLMs for Argument Mining in Education: Argument Component Identification, Classification, and Assessment
by: Favero, Lucile, et al.
Published: (2025)
by: Favero, Lucile, et al.
Published: (2025)
ToXCL: A Unified Framework for Toxic Speech Detection and Explanation
by: Hoang, Nhat M., et al.
Published: (2024)
by: Hoang, Nhat M., et al.
Published: (2024)
Few Dimensions are Enough: Fine-tuning BERT with Selected Dimensions Revealed Its Redundant Nature
by: Fukuhata, Shion, et al.
Published: (2025)
by: Fukuhata, Shion, et al.
Published: (2025)
Tox-BART: Leveraging Toxicity Attributes for Explanation Generation of Implicit Hate Speech
by: Yadav, Neemesh, et al.
Published: (2024)
by: Yadav, Neemesh, et al.
Published: (2024)
Explanation Generation for Contradiction Reconciliation with LLMs
by: Chan, Jason, et al.
Published: (2026)
by: Chan, Jason, et al.
Published: (2026)
Toxic Memes: A Survey of Computational Perspectives on the Detection and Explanation of Meme Toxicities
by: Pandiani, Delfina Sol Martinez, et al.
Published: (2024)
by: Pandiani, Delfina Sol Martinez, et al.
Published: (2024)
GloSS over Toxicity: Understanding and Mitigating Toxicity in LLMs via Global Toxic Subspace
by: Duan, Zenghao, et al.
Published: (2025)
by: Duan, Zenghao, et al.
Published: (2025)
Chinese Toxic Language Mitigation via Sentiment Polarity Consistent Rewrites
by: Wang, Xintong, et al.
Published: (2025)
by: Wang, Xintong, et al.
Published: (2025)
The Consensus Trap: Dissecting Subjectivity and the "Ground Truth" Illusion in Data Annotation
by: Munir, Sheza, et al.
Published: (2026)
by: Munir, Sheza, et al.
Published: (2026)
An LLM-Based System for Argument Mining
by: Pirozelli, Paulo, et al.
Published: (2026)
by: Pirozelli, Paulo, et al.
Published: (2026)
ArgBench: Benchmarking LLMs on Computational Argumentation Tasks
by: Ajjour, Yamen, et al.
Published: (2026)
by: Ajjour, Yamen, et al.
Published: (2026)
Argumentation for Explainable and Globally Contestable Decision Support with LLMs
by: Dejl, Adam, et al.
Published: (2026)
by: Dejl, Adam, et al.
Published: (2026)
DRIV-EX: Counterfactual Explanations for Driving LLMs
by: Cardiel, Amaia, et al.
Published: (2026)
by: Cardiel, Amaia, et al.
Published: (2026)
Can LLMs Recognize Toxicity? A Structured Investigation Framework and Toxicity Metric
by: Koh, Hyukhun, et al.
Published: (2024)
by: Koh, Hyukhun, et al.
Published: (2024)
Fearful Falcons and Angry Llamas: Emotion Category Annotations of Arguments by Humans and LLMs
by: Greschner, Lynn, et al.
Published: (2024)
by: Greschner, Lynn, et al.
Published: (2024)
Teaching LLMs Human-Like Editing of Inappropriate Argumentation via Reinforcement Learning
by: Ziegenbein, Timon, et al.
Published: (2026)
by: Ziegenbein, Timon, et al.
Published: (2026)
Is Safer Better? The Impact of Guardrails on the Argumentative Strength of LLMs in Hate Speech Countering
by: Bonaldi, Helena, et al.
Published: (2024)
by: Bonaldi, Helena, et al.
Published: (2024)
Aspect-Based Opinion Summarization with Argumentation Schemes
by: Zhou, Wendi, et al.
Published: (2025)
by: Zhou, Wendi, et al.
Published: (2025)
Improving the Reliability of LLMs: Combining CoT, RAG, Self-Consistency, and Self-Verification
by: Kumar, Adarsh, et al.
Published: (2025)
by: Kumar, Adarsh, et al.
Published: (2025)
Lost in Stories: Consistency Bugs in Long Story Generation by LLMs
by: Li, Junjie, et al.
Published: (2026)
by: Li, Junjie, et al.
Published: (2026)
Local Explanations and Self-Explanations for Assessing Faithfulness in black-box LLMs
by: Fragkathoulas, Christos, et al.
Published: (2024)
by: Fragkathoulas, Christos, et al.
Published: (2024)
Similar Items
-
Towards a Non-Ideal Methodological Framework for Responsible ML
by: Mothilal, Ramaravind Kommiya, et al.
Published: (2024) -
Reasoning About Reasoning: Towards Informed and Reflective Use of LLM Reasoning in HCI
by: Mothilal, Ramaravind Kommiya, et al.
Published: (2025) -
Talking About the Assumption in the Room
by: Mothilal, Ramaravind Kommiya, et al.
Published: (2025) -
BTPD: A Multilingual Hand-curated Dataset of Bengali Transnational Political Discourse Across Online Communities
by: Das, Dipto, et al.
Published: (2025) -
Bureaucratic Silences: What the Canadian AI Register Reveals, Omits, and Obscures
by: Das, Dipto, et al.
Published: (2026)