DetoxLLM: A Framework for Detoxification with Explanations
Fuente:
arXiv
Saved in:
| Main Authors: | Khondaker, Md Tawkat Islam, Abdul-Mageed, Muhammad, Lakshmanan, Laks V. S. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NurseLLM: The First Specialized Language Model for Nursing
by: Khondaker, Md Tawkat Islam, et al.
Published: (2025)
by: Khondaker, Md Tawkat Islam, et al.
Published: (2025)
LLM Performance Predictors are good initializers for Architecture Search
by: Jawahar, Ganesh, et al.
Published: (2023)
by: Jawahar, Ganesh, et al.
Published: (2023)
Autoregressive + Chain of Thought = Recurrent: Recurrence's Role in Language Models' Computability and a Revisit of Recurrent Transformer
by: Zhang, Xiang, et al.
Published: (2024)
by: Zhang, Xiang, et al.
Published: (2024)
UniDetox: Universal Detoxification of Large Language Models via Dataset Distillation
by: Lu, Huimin, et al.
Published: (2025)
by: Lu, Huimin, et al.
Published: (2025)
Selective Explanations
by: Paes, Lucas Monteiro, et al.
Published: (2024)
by: Paes, Lucas Monteiro, et al.
Published: (2024)
Who Gets Which Message? Auditing Demographic Bias in LLM-Generated Targeted Text
by: Islam, Tunazzina
Published: (2026)
by: Islam, Tunazzina
Published: (2026)
GRILE: A Benchmark for Grammar Reasoning and Explanation in Romanian LLMs
by: Dumitran, Adrian-Marius, et al.
Published: (2025)
by: Dumitran, Adrian-Marius, et al.
Published: (2025)
Properties and Challenges of LLM-Generated Explanations
by: Kunz, Jenny, et al.
Published: (2024)
by: Kunz, Jenny, et al.
Published: (2024)
A Multi-LLM Debiasing Framework
by: Owens, Deonna M., et al.
Published: (2024)
by: Owens, Deonna M., et al.
Published: (2024)
Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing
by: Ding, Dujian, et al.
Published: (2024)
by: Ding, Dujian, et al.
Published: (2024)
A Regularized LSTM Method for Detecting Fake News Articles
by: Camelia, Tanjina Sultana, et al.
Published: (2024)
by: Camelia, Tanjina Sultana, et al.
Published: (2024)
Exploring the Potential of the Large Language Models (LLMs) in Identifying Misleading News Headlines
by: Rony, Md Main Uddin, et al.
Published: (2024)
by: Rony, Md Main Uddin, et al.
Published: (2024)
Bangla Fake News Detection Based On Multichannel Combined CNN-LSTM
by: George, Md. Zahin Hossain, et al.
Published: (2025)
by: George, Md. Zahin Hossain, et al.
Published: (2025)
Reasoning-Based Refinement of Unsupervised Text Clusters with LLMs
by: Islam, Tunazzina
Published: (2026)
by: Islam, Tunazzina
Published: (2026)
Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations
by: Hamman, Faisal, et al.
Published: (2025)
by: Hamman, Faisal, et al.
Published: (2025)
Prompt-Counterfactual Explanations for Generative AI System Behavior
by: Goethals, Sofie, et al.
Published: (2026)
by: Goethals, Sofie, et al.
Published: (2026)
GECOBench: A Gender-Controlled Text Dataset and Benchmark for Quantifying Biases in Explanations
by: Wilming, Rick, et al.
Published: (2024)
by: Wilming, Rick, et al.
Published: (2024)
DetoxBench: Benchmarking Large Language Models for Multitask Fraud & Abuse Detection
by: Chakraborty, Joymallya, et al.
Published: (2024)
by: Chakraborty, Joymallya, et al.
Published: (2024)
Unintended Impacts of LLM Alignment on Global Representation
by: Ryan, Michael J., et al.
Published: (2024)
by: Ryan, Michael J., et al.
Published: (2024)
From Text to Talent: A Pipeline for Extracting Insights from Candidate Profiles
by: Frazzetto, Paolo, et al.
Published: (2025)
by: Frazzetto, Paolo, et al.
Published: (2025)
TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law
by: Hui, Zheng, et al.
Published: (2025)
by: Hui, Zheng, et al.
Published: (2025)
Cooperation, Competition, and Maliciousness: LLM-Stakeholders Interactive Negotiation
by: Abdelnabi, Sahar, et al.
Published: (2023)
by: Abdelnabi, Sahar, et al.
Published: (2023)
Strategic Demonstration Selection for Improved Fairness in LLM In-Context Learning
by: Hu, Jingyu, et al.
Published: (2024)
by: Hu, Jingyu, et al.
Published: (2024)
Toward LLM-Supported Automated Assessment of Critical Thinking Subskills
by: Peczuh, Marisa C., et al.
Published: (2025)
by: Peczuh, Marisa C., et al.
Published: (2025)
COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act
by: Guldimann, Philipp, et al.
Published: (2024)
by: Guldimann, Philipp, et al.
Published: (2024)
Usable XAI: 10 Strategies Towards Exploiting Explainability in the LLM Era
by: Wu, Xuansheng, et al.
Published: (2024)
by: Wu, Xuansheng, et al.
Published: (2024)
Value Drifts: Tracing Value Alignment During LLM Post-Training
by: Bhatia, Mehar, et al.
Published: (2025)
by: Bhatia, Mehar, et al.
Published: (2025)
Simulated Adoption: Decoupling Magnitude and Direction in LLM In-Context Conflict Resolution
by: Zhang, Long, et al.
Published: (2026)
by: Zhang, Long, et al.
Published: (2026)
Hypothesis Testing for Quantifying LLM-Human Misalignment in Multiple Choice Settings
by: Hong, Harbin, et al.
Published: (2025)
by: Hong, Harbin, et al.
Published: (2025)
ylmmcl at Multilingual Text Detoxification 2025: Lexicon-Guided Detoxification and Classifier-Gated Rewriting
by: Lai-Lopez, Nicole, et al.
Published: (2025)
by: Lai-Lopez, Nicole, et al.
Published: (2025)
AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts
by: Ghosh, Shaona, et al.
Published: (2024)
by: Ghosh, Shaona, et al.
Published: (2024)
Large Language Models as Urban Residents: An LLM Agent Framework for Personal Mobility Generation
by: Wang, Jiawei, et al.
Published: (2024)
by: Wang, Jiawei, et al.
Published: (2024)
AI Enabled User-Specific Cyberbullying Severity Detection with Explainability
by: Prama, Tabia Tanzin, et al.
Published: (2025)
by: Prama, Tabia Tanzin, et al.
Published: (2025)
Embracing Imperfection: Simulating Students with Diverse Cognitive Levels Using LLM-based Agents
by: Wu, Tao, et al.
Published: (2025)
by: Wu, Tao, et al.
Published: (2025)
Accept or Deny? Evaluating LLM Fairness and Performance in Loan Approval across Table-to-Text Serialization Approaches
by: Azime, Israel Abebe, et al.
Published: (2025)
by: Azime, Israel Abebe, et al.
Published: (2025)
ELMES: An Automated Framework for Evaluating Large Language Models in Educational Scenarios
by: Wei, Shou'ang, et al.
Published: (2025)
by: Wei, Shou'ang, et al.
Published: (2025)
GemDetox at TextDetox CLEF 2025: Enhancing a Massively Multilingual Model for Text Detoxification on Low-resource Languages
by: Dang, Trung Duc Anh, et al.
Published: (2025)
by: Dang, Trung Duc Anh, et al.
Published: (2025)
Reflecting in the Reflection: Integrating a Socratic Questioning Framework into Automated AI-Based Question Generation
by: Holub, Ondřej, et al.
Published: (2026)
by: Holub, Ondřej, et al.
Published: (2026)
LLM Analysis of 150+ years of German Parliamentary Debates on Migration Reveals Shift from Post-War Solidarity to Anti-Solidarity in the Last Decade
by: Kostikova, Aida, et al.
Published: (2025)
by: Kostikova, Aida, et al.
Published: (2025)
Dhoroni: Exploring Bengali Climate Change and Environmental Views with a Multi-Perspective News Dataset and Natural Language Processing
by: Wasi, Azmine Toushik, et al.
Published: (2024)
by: Wasi, Azmine Toushik, et al.
Published: (2024)
Similar Items
-
NurseLLM: The First Specialized Language Model for Nursing
by: Khondaker, Md Tawkat Islam, et al.
Published: (2025) -
LLM Performance Predictors are good initializers for Architecture Search
by: Jawahar, Ganesh, et al.
Published: (2023) -
Autoregressive + Chain of Thought = Recurrent: Recurrence's Role in Language Models' Computability and a Revisit of Recurrent Transformer
by: Zhang, Xiang, et al.
Published: (2024) -
UniDetox: Universal Detoxification of Large Language Models via Dataset Distillation
by: Lu, Huimin, et al.
Published: (2025) -
Selective Explanations
by: Paes, Lucas Monteiro, et al.
Published: (2024)