DetoxLLM: A Framework for Detoxification with Explanations
Fuente:
arXiv
Salvato in:
| Autori principali: | Khondaker, Md Tawkat Islam, Abdul-Mageed, Muhammad, Lakshmanan, Laks V. S. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
NurseLLM: The First Specialized Language Model for Nursing
di: Khondaker, Md Tawkat Islam, et al.
Pubblicazione: (2025)
di: Khondaker, Md Tawkat Islam, et al.
Pubblicazione: (2025)
LLM Performance Predictors are good initializers for Architecture Search
di: Jawahar, Ganesh, et al.
Pubblicazione: (2023)
di: Jawahar, Ganesh, et al.
Pubblicazione: (2023)
Autoregressive + Chain of Thought = Recurrent: Recurrence's Role in Language Models' Computability and a Revisit of Recurrent Transformer
di: Zhang, Xiang, et al.
Pubblicazione: (2024)
di: Zhang, Xiang, et al.
Pubblicazione: (2024)
UniDetox: Universal Detoxification of Large Language Models via Dataset Distillation
di: Lu, Huimin, et al.
Pubblicazione: (2025)
di: Lu, Huimin, et al.
Pubblicazione: (2025)
Selective Explanations
di: Paes, Lucas Monteiro, et al.
Pubblicazione: (2024)
di: Paes, Lucas Monteiro, et al.
Pubblicazione: (2024)
Who Gets Which Message? Auditing Demographic Bias in LLM-Generated Targeted Text
di: Islam, Tunazzina
Pubblicazione: (2026)
di: Islam, Tunazzina
Pubblicazione: (2026)
GRILE: A Benchmark for Grammar Reasoning and Explanation in Romanian LLMs
di: Dumitran, Adrian-Marius, et al.
Pubblicazione: (2025)
di: Dumitran, Adrian-Marius, et al.
Pubblicazione: (2025)
Properties and Challenges of LLM-Generated Explanations
di: Kunz, Jenny, et al.
Pubblicazione: (2024)
di: Kunz, Jenny, et al.
Pubblicazione: (2024)
A Multi-LLM Debiasing Framework
di: Owens, Deonna M., et al.
Pubblicazione: (2024)
di: Owens, Deonna M., et al.
Pubblicazione: (2024)
Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing
di: Ding, Dujian, et al.
Pubblicazione: (2024)
di: Ding, Dujian, et al.
Pubblicazione: (2024)
A Regularized LSTM Method for Detecting Fake News Articles
di: Camelia, Tanjina Sultana, et al.
Pubblicazione: (2024)
di: Camelia, Tanjina Sultana, et al.
Pubblicazione: (2024)
Exploring the Potential of the Large Language Models (LLMs) in Identifying Misleading News Headlines
di: Rony, Md Main Uddin, et al.
Pubblicazione: (2024)
di: Rony, Md Main Uddin, et al.
Pubblicazione: (2024)
Bangla Fake News Detection Based On Multichannel Combined CNN-LSTM
di: George, Md. Zahin Hossain, et al.
Pubblicazione: (2025)
di: George, Md. Zahin Hossain, et al.
Pubblicazione: (2025)
Reasoning-Based Refinement of Unsupervised Text Clusters with LLMs
di: Islam, Tunazzina
Pubblicazione: (2026)
di: Islam, Tunazzina
Pubblicazione: (2026)
Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations
di: Hamman, Faisal, et al.
Pubblicazione: (2025)
di: Hamman, Faisal, et al.
Pubblicazione: (2025)
Prompt-Counterfactual Explanations for Generative AI System Behavior
di: Goethals, Sofie, et al.
Pubblicazione: (2026)
di: Goethals, Sofie, et al.
Pubblicazione: (2026)
GECOBench: A Gender-Controlled Text Dataset and Benchmark for Quantifying Biases in Explanations
di: Wilming, Rick, et al.
Pubblicazione: (2024)
di: Wilming, Rick, et al.
Pubblicazione: (2024)
DetoxBench: Benchmarking Large Language Models for Multitask Fraud & Abuse Detection
di: Chakraborty, Joymallya, et al.
Pubblicazione: (2024)
di: Chakraborty, Joymallya, et al.
Pubblicazione: (2024)
Unintended Impacts of LLM Alignment on Global Representation
di: Ryan, Michael J., et al.
Pubblicazione: (2024)
di: Ryan, Michael J., et al.
Pubblicazione: (2024)
From Text to Talent: A Pipeline for Extracting Insights from Candidate Profiles
di: Frazzetto, Paolo, et al.
Pubblicazione: (2025)
di: Frazzetto, Paolo, et al.
Pubblicazione: (2025)
TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law
di: Hui, Zheng, et al.
Pubblicazione: (2025)
di: Hui, Zheng, et al.
Pubblicazione: (2025)
Cooperation, Competition, and Maliciousness: LLM-Stakeholders Interactive Negotiation
di: Abdelnabi, Sahar, et al.
Pubblicazione: (2023)
di: Abdelnabi, Sahar, et al.
Pubblicazione: (2023)
Strategic Demonstration Selection for Improved Fairness in LLM In-Context Learning
di: Hu, Jingyu, et al.
Pubblicazione: (2024)
di: Hu, Jingyu, et al.
Pubblicazione: (2024)
Toward LLM-Supported Automated Assessment of Critical Thinking Subskills
di: Peczuh, Marisa C., et al.
Pubblicazione: (2025)
di: Peczuh, Marisa C., et al.
Pubblicazione: (2025)
COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act
di: Guldimann, Philipp, et al.
Pubblicazione: (2024)
di: Guldimann, Philipp, et al.
Pubblicazione: (2024)
Usable XAI: 10 Strategies Towards Exploiting Explainability in the LLM Era
di: Wu, Xuansheng, et al.
Pubblicazione: (2024)
di: Wu, Xuansheng, et al.
Pubblicazione: (2024)
Value Drifts: Tracing Value Alignment During LLM Post-Training
di: Bhatia, Mehar, et al.
Pubblicazione: (2025)
di: Bhatia, Mehar, et al.
Pubblicazione: (2025)
Simulated Adoption: Decoupling Magnitude and Direction in LLM In-Context Conflict Resolution
di: Zhang, Long, et al.
Pubblicazione: (2026)
di: Zhang, Long, et al.
Pubblicazione: (2026)
Hypothesis Testing for Quantifying LLM-Human Misalignment in Multiple Choice Settings
di: Hong, Harbin, et al.
Pubblicazione: (2025)
di: Hong, Harbin, et al.
Pubblicazione: (2025)
ylmmcl at Multilingual Text Detoxification 2025: Lexicon-Guided Detoxification and Classifier-Gated Rewriting
di: Lai-Lopez, Nicole, et al.
Pubblicazione: (2025)
di: Lai-Lopez, Nicole, et al.
Pubblicazione: (2025)
AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts
di: Ghosh, Shaona, et al.
Pubblicazione: (2024)
di: Ghosh, Shaona, et al.
Pubblicazione: (2024)
Large Language Models as Urban Residents: An LLM Agent Framework for Personal Mobility Generation
di: Wang, Jiawei, et al.
Pubblicazione: (2024)
di: Wang, Jiawei, et al.
Pubblicazione: (2024)
AI Enabled User-Specific Cyberbullying Severity Detection with Explainability
di: Prama, Tabia Tanzin, et al.
Pubblicazione: (2025)
di: Prama, Tabia Tanzin, et al.
Pubblicazione: (2025)
Embracing Imperfection: Simulating Students with Diverse Cognitive Levels Using LLM-based Agents
di: Wu, Tao, et al.
Pubblicazione: (2025)
di: Wu, Tao, et al.
Pubblicazione: (2025)
Accept or Deny? Evaluating LLM Fairness and Performance in Loan Approval across Table-to-Text Serialization Approaches
di: Azime, Israel Abebe, et al.
Pubblicazione: (2025)
di: Azime, Israel Abebe, et al.
Pubblicazione: (2025)
ELMES: An Automated Framework for Evaluating Large Language Models in Educational Scenarios
di: Wei, Shou'ang, et al.
Pubblicazione: (2025)
di: Wei, Shou'ang, et al.
Pubblicazione: (2025)
GemDetox at TextDetox CLEF 2025: Enhancing a Massively Multilingual Model for Text Detoxification on Low-resource Languages
di: Dang, Trung Duc Anh, et al.
Pubblicazione: (2025)
di: Dang, Trung Duc Anh, et al.
Pubblicazione: (2025)
Reflecting in the Reflection: Integrating a Socratic Questioning Framework into Automated AI-Based Question Generation
di: Holub, Ondřej, et al.
Pubblicazione: (2026)
di: Holub, Ondřej, et al.
Pubblicazione: (2026)
LLM Analysis of 150+ years of German Parliamentary Debates on Migration Reveals Shift from Post-War Solidarity to Anti-Solidarity in the Last Decade
di: Kostikova, Aida, et al.
Pubblicazione: (2025)
di: Kostikova, Aida, et al.
Pubblicazione: (2025)
Dhoroni: Exploring Bengali Climate Change and Environmental Views with a Multi-Perspective News Dataset and Natural Language Processing
di: Wasi, Azmine Toushik, et al.
Pubblicazione: (2024)
di: Wasi, Azmine Toushik, et al.
Pubblicazione: (2024)
Documenti analoghi
-
NurseLLM: The First Specialized Language Model for Nursing
di: Khondaker, Md Tawkat Islam, et al.
Pubblicazione: (2025) -
LLM Performance Predictors are good initializers for Architecture Search
di: Jawahar, Ganesh, et al.
Pubblicazione: (2023) -
Autoregressive + Chain of Thought = Recurrent: Recurrence's Role in Language Models' Computability and a Revisit of Recurrent Transformer
di: Zhang, Xiang, et al.
Pubblicazione: (2024) -
UniDetox: Universal Detoxification of Large Language Models via Dataset Distillation
di: Lu, Huimin, et al.
Pubblicazione: (2025) -
Selective Explanations
di: Paes, Lucas Monteiro, et al.
Pubblicazione: (2024)