Adaptive Detoxification: Safeguarding General Capabilities of LLMs through Toxicity-Aware Knowledge Editing
Fuente:
arXiv
Salvato in:
| Autori principali: | Lu, Yifan, Li, Jing, Zhou, Yigeng, Zhang, Yihui, Wang, Wenya, Li, Xiucheng, Zhang, Meishan, Liu, Fangming, Yu, Jun, Zhang, Min |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Knowledge Editing with Dynamic Knowledge Graphs for Multi-Hop Question Answering
di: Lu, Yifan, et al.
Pubblicazione: (2024)
di: Lu, Yifan, et al.
Pubblicazione: (2024)
Mitigating Context-Memory Conflicts in LLMs through Dynamic Cognitive Reconciliation Decoding
di: Zhou, Yigeng, et al.
Pubblicazione: (2026)
di: Zhou, Yigeng, et al.
Pubblicazione: (2026)
Multi-objective Large Language Model Alignment with Hierarchical Experts
di: Li, Zhuo, et al.
Pubblicazione: (2025)
di: Li, Zhuo, et al.
Pubblicazione: (2025)
Team-Based Self-Play With Dual Adaptive Weighting for Fine-Tuning LLMs
di: Li, Wu, et al.
Pubblicazione: (2026)
di: Li, Wu, et al.
Pubblicazione: (2026)
Jailbreak-R1: Exploring the Jailbreak Capabilities of LLMs via Reinforcement Learning
di: Guo, Weiyang, et al.
Pubblicazione: (2025)
di: Guo, Weiyang, et al.
Pubblicazione: (2025)
Knowledge Fusion of Large Language Models Via Modular SkillPacks
di: Du, Guodong, et al.
Pubblicazione: (2025)
di: Du, Guodong, et al.
Pubblicazione: (2025)
On the Robustness of Knowledge Editing for Detoxification
di: Dong, Ming, et al.
Pubblicazione: (2026)
di: Dong, Ming, et al.
Pubblicazione: (2026)
Modality-Decoupled Online Recursive Editing
di: Li, Siyuan, et al.
Pubblicazione: (2026)
di: Li, Siyuan, et al.
Pubblicazione: (2026)
LLMs Can Also Do Well! Breaking Barriers in Semantic Role Labeling via Large Language Models
di: Li, Xinxin, et al.
Pubblicazione: (2025)
di: Li, Xinxin, et al.
Pubblicazione: (2025)
MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming
di: Guo, Weiyang, et al.
Pubblicazione: (2025)
di: Guo, Weiyang, et al.
Pubblicazione: (2025)
Multi-Modality Expansion and Retention for LLMs through Parameter Merging and Decoupling
di: Li, Junlin, et al.
Pubblicazione: (2025)
di: Li, Junlin, et al.
Pubblicazione: (2025)
MobiEdit: Resource-efficient Knowledge Editing for Personalized On-device LLMs
di: Lu, Zhenyan, et al.
Pubblicazione: (2025)
di: Lu, Zhenyan, et al.
Pubblicazione: (2025)
Function-to-Style Guidance of LLMs for Code Translation
di: Zhang, Longhui, et al.
Pubblicazione: (2025)
di: Zhang, Longhui, et al.
Pubblicazione: (2025)
Knowledgeable Preference Alignment for LLMs in Domain-specific Question Answering
di: Zhang, Yichi, et al.
Pubblicazione: (2023)
di: Zhang, Yichi, et al.
Pubblicazione: (2023)
Uncovering the Fragility of Trustworthy LLMs through Chinese Textual Ambiguity
di: Wu, Xinwei, et al.
Pubblicazione: (2025)
di: Wu, Xinwei, et al.
Pubblicazione: (2025)
DEVAL: A Framework for Evaluating and Improving the Derivation Capability of Large Language Models
di: Li, Yifan, et al.
Pubblicazione: (2025)
di: Li, Yifan, et al.
Pubblicazione: (2025)
CoreGuard: Safeguarding Foundational Capabilities of LLMs Against Model Stealing in Edge Deployment
di: Li, Qinfeng, et al.
Pubblicazione: (2024)
di: Li, Qinfeng, et al.
Pubblicazione: (2024)
Contrastive Learning on LLM Back Generation Treebank for Cross-domain Constituency Parsing
di: Guo, Peiming, et al.
Pubblicazione: (2025)
di: Guo, Peiming, et al.
Pubblicazione: (2025)
An Adaptive Finite Element Method Based on Generalized Barycentric Coordinates
di: Zhou, Yihui, et al.
Pubblicazione: (2026)
di: Zhou, Yihui, et al.
Pubblicazione: (2026)
BrainECHO: Semantic Brain Signal Decoding through Vector-Quantized Spectrogram Reconstruction for Whisper-Enhanced Text Generation
di: Li, Jilong, et al.
Pubblicazione: (2024)
di: Li, Jilong, et al.
Pubblicazione: (2024)
DAPI: Domain Adaptive Toxicity Probe Vector Intervention for Fine-Grained Detoxification
di: Hyeonsu, Cho, et al.
Pubblicazione: (2025)
di: Hyeonsu, Cho, et al.
Pubblicazione: (2025)
Beyond Chunking: Discourse-Aware Hierarchical Retrieval for Long Document Question Answering
di: Chen, Huiyao, et al.
Pubblicazione: (2025)
di: Chen, Huiyao, et al.
Pubblicazione: (2025)
Safeguarding Text-to-Image Generative Models Against Unauthorized Knowledge Distillation
di: Gao, Yilan, et al.
Pubblicazione: (2026)
di: Gao, Yilan, et al.
Pubblicazione: (2026)
Text Detoxification: Data Efficiency, Semantic Preservation and Model Generalization
di: Yu, Jing, et al.
Pubblicazione: (2025)
di: Yu, Jing, et al.
Pubblicazione: (2025)
Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs
di: Fan, Zhaoyu, et al.
Pubblicazione: (2025)
di: Fan, Zhaoyu, et al.
Pubblicazione: (2025)
Chlorella Vulgaris‐Inspired Versatile Theranostic Nanoparticles for Specific Recognition and Detoxification to Copper (II) In Vitro and In Vivo
di: Xu‐Wei Qi, et al.
Pubblicazione: (2024)
di: Xu‐Wei Qi, et al.
Pubblicazione: (2024)
PropGuard: Safeguarding LLM-MAS via Propagation-Aware Exploration and Remediation
di: Yan, Bingyu, et al.
Pubblicazione: (2026)
di: Yan, Bingyu, et al.
Pubblicazione: (2026)
ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs
di: Cui, Shiyao, et al.
Pubblicazione: (2025)
di: Cui, Shiyao, et al.
Pubblicazione: (2025)
Chinese Sequence Labeling with Semi-Supervised Boundary-Aware Language Model Pre-training
di: Zhang, Longhui, et al.
Pubblicazione: (2024)
di: Zhang, Longhui, et al.
Pubblicazione: (2024)
APAO: Adaptive Prefix-Aware Optimization for Generative Recommendation
di: Yu, Yuanqing, et al.
Pubblicazione: (2026)
di: Yu, Yuanqing, et al.
Pubblicazione: (2026)
Enhancing the General Agent Capabilities of Low-Parameter LLMs through Tuning and Multi-Branch Reasoning
di: Zhou, Qinhao, et al.
Pubblicazione: (2024)
di: Zhou, Qinhao, et al.
Pubblicazione: (2024)
Revealing and Mitigating Over-Attention in Knowledge Editing
di: Wang, Pinzheng, et al.
Pubblicazione: (2025)
di: Wang, Pinzheng, et al.
Pubblicazione: (2025)
GME: Improving Universal Multimodal Retrieval by Multimodal LLMs
di: Zhang, Xin, et al.
Pubblicazione: (2024)
di: Zhang, Xin, et al.
Pubblicazione: (2024)
CMD: a framework for Context-aware Model self-Detoxification
di: Tang, Zecheng, et al.
Pubblicazione: (2023)
di: Tang, Zecheng, et al.
Pubblicazione: (2023)
KGEdit: Ambiguity-Aware Knowledge Graphs for Training-Free Precise Video Generation and Editing
di: Cai, Mingshu, et al.
Pubblicazione: (2026)
di: Cai, Mingshu, et al.
Pubblicazione: (2026)
Dynamic Long Context Reasoning over Compressed Memory via End-to-End Reinforcement Learning
di: Chen, Zhuoen, et al.
Pubblicazione: (2026)
di: Chen, Zhuoen, et al.
Pubblicazione: (2026)
Echoes as Anchors: Probabilistic Costs and Attention Refocusing in LLM Reasoning
di: Hao, Zhuoyuan, et al.
Pubblicazione: (2026)
di: Hao, Zhuoyuan, et al.
Pubblicazione: (2026)
A Policy Report Evaluating the National Assessment Program for Literacy and Numeracy (Naplan) Reform in Australia: The Impacts of High Stakes Assessment on Students
di: Zhang, Wenya
Pubblicazione: (2024)
di: Zhang, Wenya
Pubblicazione: (2024)
ToolOmni: Enabling Open-World Tool Use via Agentic learning with Proactive Retrieval and Grounded Execution
di: Huang, Shouzheng, et al.
Pubblicazione: (2026)
di: Huang, Shouzheng, et al.
Pubblicazione: (2026)
In-Context Learning for Few-Shot Nested Named Entity Recognition
di: Zhang, Meishan, et al.
Pubblicazione: (2024)
di: Zhang, Meishan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Knowledge Editing with Dynamic Knowledge Graphs for Multi-Hop Question Answering
di: Lu, Yifan, et al.
Pubblicazione: (2024) -
Mitigating Context-Memory Conflicts in LLMs through Dynamic Cognitive Reconciliation Decoding
di: Zhou, Yigeng, et al.
Pubblicazione: (2026) -
Multi-objective Large Language Model Alignment with Hierarchical Experts
di: Li, Zhuo, et al.
Pubblicazione: (2025) -
Team-Based Self-Play With Dual Adaptive Weighting for Fine-Tuning LLMs
di: Li, Wu, et al.
Pubblicazione: (2026) -
Jailbreak-R1: Exploring the Jailbreak Capabilities of LLMs via Reinforcement Learning
di: Guo, Weiyang, et al.
Pubblicazione: (2025)