Detoxification for LLM: From Dataset Itself
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shao, Wei, Wang, Yihang, Zhu, Gaoyu, Cheng, Ziqiang, Yu, Lei, Guo, Jiafeng, Cheng, Xueqi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How Knowledge Popularity Influences and Enhances LLM Knowledge Boundary Perception
von: Ni, Shiyu, et al.
Veröffentlicht: (2025)
von: Ni, Shiyu, et al.
Veröffentlicht: (2025)
Evaluating Implicit Bias in Large Language Models by Attacking From a Psychometric Perspective
von: Wen, Yuchen, et al.
Veröffentlicht: (2024)
von: Wen, Yuchen, et al.
Veröffentlicht: (2024)
MVAM: Multi-View Attention Method for Fine-grained Image-Text Matching
von: Cui, Wanqing, et al.
Veröffentlicht: (2024)
von: Cui, Wanqing, et al.
Veröffentlicht: (2024)
Towards Fully Exploiting LLM Internal States to Enhance Knowledge Boundary Perception
von: Ni, Shiyu, et al.
Veröffentlicht: (2025)
von: Ni, Shiyu, et al.
Veröffentlicht: (2025)
How Do LLM-Generated Texts Impact Term-Based Retrieval Models?
von: Huang, Wei, et al.
Veröffentlicht: (2025)
von: Huang, Wei, et al.
Veröffentlicht: (2025)
MORE: Multi-mOdal REtrieval Augmented Generative Commonsense Reasoning
von: Cui, Wanqing, et al.
Veröffentlicht: (2024)
von: Cui, Wanqing, et al.
Veröffentlicht: (2024)
When Do LLMs Need Retrieval Augmentation? Mitigating LLMs' Overconfidence Helps Retrieval Augmentation
von: Ni, Shiyu, et al.
Veröffentlicht: (2024)
von: Ni, Shiyu, et al.
Veröffentlicht: (2024)
QUITO-X: A New Perspective on Context Compression from the Information Bottleneck Theory
von: Wang, Yihang, et al.
Veröffentlicht: (2024)
von: Wang, Yihang, et al.
Veröffentlicht: (2024)
Estimating Commonsense Plausibility through Semantic Shifts
von: Cui, Wanqing, et al.
Veröffentlicht: (2025)
von: Cui, Wanqing, et al.
Veröffentlicht: (2025)
A Unified Causal View of Instruction Tuning
von: Chen, Lu, et al.
Veröffentlicht: (2024)
von: Chen, Lu, et al.
Veröffentlicht: (2024)
Iterative Structured Pruning for Large Language Models with Multi-Domain Calibration
von: Wu, Guangxin, et al.
Veröffentlicht: (2026)
von: Wu, Guangxin, et al.
Veröffentlicht: (2026)
LINKAGE: Listwise Ranking among Varied-Quality References for Non-Factoid QA Evaluation via LLMs
von: Yang, Sihui, et al.
Veröffentlicht: (2024)
von: Yang, Sihui, et al.
Veröffentlicht: (2024)
Controlling Risk of Retrieval-augmented Generation: A Counterfactual Prompting Framework
von: Chen, Lu, et al.
Veröffentlicht: (2024)
von: Chen, Lu, et al.
Veröffentlicht: (2024)
Class-Incremental Few-Shot Event Detection
von: Zhao, Kailin, et al.
Veröffentlicht: (2024)
von: Zhao, Kailin, et al.
Veröffentlicht: (2024)
An Iterative Utility Judgment Framework Inspired by Philosophical Relevance via LLMs
von: Zhang, Hengran, et al.
Veröffentlicht: (2024)
von: Zhang, Hengran, et al.
Veröffentlicht: (2024)
LLM-Specific Utility: A New Perspective for Retrieval-Augmented Generation
von: Zhang, Hengran, et al.
Veröffentlicht: (2025)
von: Zhang, Hengran, et al.
Veröffentlicht: (2025)
Qsnail: A Questionnaire Dataset for Sequential Question Generation
von: Lei, Yan, et al.
Veröffentlicht: (2024)
von: Lei, Yan, et al.
Veröffentlicht: (2024)
EGAD: Entropy-Guided Adaptive Distillation for Token-Level Knowledge Transfer
von: Zhang, Hao, et al.
Veröffentlicht: (2026)
von: Zhang, Hao, et al.
Veröffentlicht: (2026)
MI-PRUN: Optimize Large Language Model Pruning via Mutual Information
von: Zhang, Hao, et al.
Veröffentlicht: (2026)
von: Zhang, Hao, et al.
Veröffentlicht: (2026)
SelfCP: Compressing Over-Limit Prompt via the Frozen Large Language Model Itself
von: Gao, Jun, et al.
Veröffentlicht: (2024)
von: Gao, Jun, et al.
Veröffentlicht: (2024)
Self-Guard: Empower the LLM to Safeguard Itself
von: Wang, Zezhong, et al.
Veröffentlicht: (2023)
von: Wang, Zezhong, et al.
Veröffentlicht: (2023)
StruProKGR: A Structural and Probabilistic Framework for Sparse Knowledge Graph Reasoning
von: Guo, Yucan, et al.
Veröffentlicht: (2025)
von: Guo, Yucan, et al.
Veröffentlicht: (2025)
LLM in the Loop: Creating the ParaDeHate Dataset for Hate Speech Detoxification
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
G2S: A General-to-Specific Learning Framework for Temporal Knowledge Graph Forecasting with Large Language Models
von: Bai, Long, et al.
Veröffentlicht: (2025)
von: Bai, Long, et al.
Veröffentlicht: (2025)
Towards Event Extraction with Massive Types: LLM-based Collaborative Annotation and Partitioning Extraction
von: Liu, Wenxuan, et al.
Veröffentlicht: (2025)
von: Liu, Wenxuan, et al.
Veröffentlicht: (2025)
Towards Robust Universal Information Extraction: Benchmark, Evaluation, and Solution
von: Zhu, Jizhao, et al.
Veröffentlicht: (2025)
von: Zhu, Jizhao, et al.
Veröffentlicht: (2025)
Utility-Focused LLM Annotation for Retrieval and Retrieval-Augmented Generation
von: Zhang, Hengran, et al.
Veröffentlicht: (2025)
von: Zhang, Hengran, et al.
Veröffentlicht: (2025)
Annotation-Efficient Universal Honesty Alignment
von: Ni, Shiyu, et al.
Veröffentlicht: (2025)
von: Ni, Shiyu, et al.
Veröffentlicht: (2025)
Pretraining Data Detection for Large Language Models: A Divergence-based Calibration Method
von: Zhang, Weichao, et al.
Veröffentlicht: (2024)
von: Zhang, Weichao, et al.
Veröffentlicht: (2024)
An In-Context Schema Understanding Method for Knowledge Base Question Answering
von: Liu, Yantao, et al.
Veröffentlicht: (2023)
von: Liu, Yantao, et al.
Veröffentlicht: (2023)
Robust Neural Information Retrieval: An Adversarial and Out-of-distribution Perspective
von: Liu, Yu-An, et al.
Veröffentlicht: (2024)
von: Liu, Yu-An, et al.
Veröffentlicht: (2024)
RouteRAG: Efficient Retrieval-Augmented Generation from Text and Graph via Reinforcement Learning
von: Guo, Yucan, et al.
Veröffentlicht: (2025)
von: Guo, Yucan, et al.
Veröffentlicht: (2025)
QUITO: Accelerating Long-Context Reasoning through Query-Guided Context Compression
von: Wang, Wenshan, et al.
Veröffentlicht: (2024)
von: Wang, Wenshan, et al.
Veröffentlicht: (2024)
DetoxLLM: A Framework for Detoxification with Explanations
von: Khondaker, Md Tawkat Islam, et al.
Veröffentlicht: (2024)
von: Khondaker, Md Tawkat Islam, et al.
Veröffentlicht: (2024)
Prism-$Δ$: Differential Subspace Steering for Prompt Highlighting in Large Language Models
von: Ge, Yuyao, et al.
Veröffentlicht: (2026)
von: Ge, Yuyao, et al.
Veröffentlicht: (2026)
CorpusBrain++: A Continual Generative Pre-Training Framework for Knowledge-Intensive Language Tasks
von: Guo, Jiafeng, et al.
Veröffentlicht: (2024)
von: Guo, Jiafeng, et al.
Veröffentlicht: (2024)
Text Detoxification: Data Efficiency, Semantic Preservation and Model Generalization
von: Yu, Jing, et al.
Veröffentlicht: (2025)
von: Yu, Jing, et al.
Veröffentlicht: (2025)
Thinking Forward and Backward: Multi-Objective Reinforcement Learning for Retrieval-Augmented Reasoning
von: Wei, Wenda, et al.
Veröffentlicht: (2025)
von: Wei, Wenda, et al.
Veröffentlicht: (2025)
Bagging-Based Model Merging for Robust General Text Embeddings
von: Zhang, Hengran, et al.
Veröffentlicht: (2026)
von: Zhang, Hengran, et al.
Veröffentlicht: (2026)
Distilling a Small Utility-Based Passage Selector to Enhance Retrieval-Augmented Generation
von: Zhang, Hengran, et al.
Veröffentlicht: (2025)
von: Zhang, Hengran, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
How Knowledge Popularity Influences and Enhances LLM Knowledge Boundary Perception
von: Ni, Shiyu, et al.
Veröffentlicht: (2025) -
Evaluating Implicit Bias in Large Language Models by Attacking From a Psychometric Perspective
von: Wen, Yuchen, et al.
Veröffentlicht: (2024) -
MVAM: Multi-View Attention Method for Fine-grained Image-Text Matching
von: Cui, Wanqing, et al.
Veröffentlicht: (2024) -
Towards Fully Exploiting LLM Internal States to Enhance Knowledge Boundary Perception
von: Ni, Shiyu, et al.
Veröffentlicht: (2025) -
How Do LLM-Generated Texts Impact Term-Based Retrieval Models?
von: Huang, Wei, et al.
Veröffentlicht: (2025)