Detoxifying Large Language Models via Autoregressive Reward Guided Representation Editing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xiao, Yisong, Liu, Aishan, Liang, Siyuan, Ying, Zonghao, Liu, Xianglong, Tao, Dacheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
Fairness Mediator: Neutralize Stereotype Associations to Mitigate Bias in Large Language Models
von: Xiao, Yisong, et al.
Veröffentlicht: (2025)
von: Xiao, Yisong, et al.
Veröffentlicht: (2025)
Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
Unveiling the Safety of GPT-4o: An Empirical Study using Jailbreak Attacks
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
Uncovering Strategic Egoism Behaviors in Large Language Models
von: Zhang, Yaoyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Yaoyuan, et al.
Veröffentlicht: (2025)
CogMorph: Cognitive Morphing Attacks for Text-to-Image Models
von: Jing, Zonglei, et al.
Veröffentlicht: (2025)
von: Jing, Zonglei, et al.
Veröffentlicht: (2025)
Detoxifying Large Language Models via Knowledge Editing
von: Wang, Mengru, et al.
Veröffentlicht: (2024)
von: Wang, Mengru, et al.
Veröffentlicht: (2024)
Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
GenderBias-\emph{VL}: Benchmarking Gender Bias in Vision Language Models via Counterfactual Probing
von: Xiao, Yisong, et al.
Veröffentlicht: (2024)
von: Xiao, Yisong, et al.
Veröffentlicht: (2024)
LanEvil: Benchmarking the Robustness of Lane Detection to Environmental Illusions
von: Zhang, Tianyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Tianyuan, et al.
Veröffentlicht: (2024)
Revisiting Knowledge Distillation for Autoregressive Language Models
von: Zhong, Qihuang, et al.
Veröffentlicht: (2024)
von: Zhong, Qihuang, et al.
Veröffentlicht: (2024)
Exploring and Enhancing the Transfer of Distribution in Knowledge Distillation for Autoregressive Language Models
von: Rao, Jun, et al.
Veröffentlicht: (2024)
von: Rao, Jun, et al.
Veröffentlicht: (2024)
Manipulating Multimodal Agents via Cross-Modal Prompt Injection
von: Wang, Le, et al.
Veröffentlicht: (2025)
von: Wang, Le, et al.
Veröffentlicht: (2025)
Learning from Imperfect Data: Towards Efficient Knowledge Distillation of Autoregressive Language Models for Text-to-SQL
von: Zhong, Qihuang, et al.
Veröffentlicht: (2024)
von: Zhong, Qihuang, et al.
Veröffentlicht: (2024)
Black-Box Adversarial Attack on Vision Language Models for Autonomous Driving
von: Wang, Lu, et al.
Veröffentlicht: (2025)
von: Wang, Lu, et al.
Veröffentlicht: (2025)
Large Language Models can be Strong Self-Detoxifiers
von: Ko, Ching-Yun, et al.
Veröffentlicht: (2024)
von: Ko, Ching-Yun, et al.
Veröffentlicht: (2024)
Towards Reliable Medical LLMs: Benchmarking and Enhancing Confidence Estimation of Large Language Models in Medical Consultation
von: Ren, Zhiyao, et al.
Veröffentlicht: (2026)
von: Ren, Zhiyao, et al.
Veröffentlicht: (2026)
NoVo: Norm Voting off Hallucinations with Attention Heads in Large Language Models
von: Ho, Zheng Yi, et al.
Veröffentlicht: (2024)
von: Ho, Zheng Yi, et al.
Veröffentlicht: (2024)
LLMCBench: Benchmarking Large Language Model Compression for Efficient Deployment
von: Yang, Ge, et al.
Veröffentlicht: (2024)
von: Yang, Ge, et al.
Veröffentlicht: (2024)
Contrastive Perplexity for Controlled Generation: An Application in Detoxifying Large Language Models
von: Klein, Tassilo, et al.
Veröffentlicht: (2024)
von: Klein, Tassilo, et al.
Veröffentlicht: (2024)
Disentangling Knowledge Representations for Large Language Model Editing
von: Zhang, Mengqi, et al.
Veröffentlicht: (2025)
von: Zhang, Mengqi, et al.
Veröffentlicht: (2025)
Robust Knowledge Editing via Explicit Reasoning Chains for Distractor-Resilient Multi-Hop QA
von: Wu, Yuchen, et al.
Veröffentlicht: (2025)
von: Wu, Yuchen, et al.
Veröffentlicht: (2025)
CodeSimpleQA: Scaling Factuality in Code Large Language Models
von: Yang, Jian, et al.
Veröffentlicht: (2025)
von: Yang, Jian, et al.
Veröffentlicht: (2025)
Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning
von: Park, Sungjin, et al.
Veröffentlicht: (2024)
von: Park, Sungjin, et al.
Veröffentlicht: (2024)
ROSE Doesn't Do That: Boosting the Safety of Instruction-Tuned Large Language Models with Reverse Prompt Contrastive Decoding
von: Zhong, Qihuang, et al.
Veröffentlicht: (2024)
von: Zhong, Qihuang, et al.
Veröffentlicht: (2024)
FourierSampler: Unlocking Non-Autoregressive Potential in Diffusion Language Models via Frequency-Guided Generation
von: He, Siyang, et al.
Veröffentlicht: (2026)
von: He, Siyang, et al.
Veröffentlicht: (2026)
LLMC: Benchmarking Large Language Model Quantization with a Versatile Compression Toolkit
von: Gong, Ruihao, et al.
Veröffentlicht: (2024)
von: Gong, Ruihao, et al.
Veröffentlicht: (2024)
Revisiting Catastrophic Forgetting in Large Language Model Tuning
von: Li, Hongyu, et al.
Veröffentlicht: (2024)
von: Li, Hongyu, et al.
Veröffentlicht: (2024)
Diversifying the Mixture-of-Experts Representation for Language Models with Orthogonal Optimizer
von: Liu, Boan, et al.
Veröffentlicht: (2023)
von: Liu, Boan, et al.
Veröffentlicht: (2023)
Genshin: General Shield for Natural Language Processing with Large Language Models
von: Peng, Xiao, et al.
Veröffentlicht: (2024)
von: Peng, Xiao, et al.
Veröffentlicht: (2024)
BDefects4NN: A Backdoor Defect Database for Controlled Localization Studies in Neural Networks
von: Xiao, Yisong, et al.
Veröffentlicht: (2024)
von: Xiao, Yisong, et al.
Veröffentlicht: (2024)
GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment
von: Xu, Yuancheng, et al.
Veröffentlicht: (2024)
von: Xu, Yuancheng, et al.
Veröffentlicht: (2024)
Reason-KE++: Aligning the Process, Not Just the Outcome, for Faithful LLM Knowledge Editing
von: Wu, Yuchen, et al.
Veröffentlicht: (2025)
von: Wu, Yuchen, et al.
Veröffentlicht: (2025)
Revisiting Backdoor Attacks against Large Vision-Language Models from Domain Shift
von: Liang, Siyuan, et al.
Veröffentlicht: (2024)
von: Liang, Siyuan, et al.
Veröffentlicht: (2024)
Robust and Scalable Model Editing for Large Language Models
von: Chen, Yingfa, et al.
Veröffentlicht: (2024)
von: Chen, Yingfa, et al.
Veröffentlicht: (2024)
EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing
von: Wu, Keming, et al.
Veröffentlicht: (2025)
von: Wu, Keming, et al.
Veröffentlicht: (2025)
ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
von: Zhang, Dan, et al.
Veröffentlicht: (2024)
von: Zhang, Dan, et al.
Veröffentlicht: (2024)
Latent Imitator: Generating Natural Individual Discriminatory Instances for Black-Box Fairness Testing
von: Xiao, Yisong, et al.
Veröffentlicht: (2023)
von: Xiao, Yisong, et al.
Veröffentlicht: (2023)
Neuron-Level Sequential Editing for Large Language Models
von: Jiang, Houcheng, et al.
Veröffentlicht: (2024)
von: Jiang, Houcheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
von: Ying, Zonghao, et al.
Veröffentlicht: (2025) -
Fairness Mediator: Neutralize Stereotype Associations to Mitigate Bias in Large Language Models
von: Xiao, Yisong, et al.
Veröffentlicht: (2025) -
Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt
von: Ying, Zonghao, et al.
Veröffentlicht: (2024) -
Unveiling the Safety of GPT-4o: An Empirical Study using Jailbreak Attacks
von: Ying, Zonghao, et al.
Veröffentlicht: (2024) -
Uncovering Strategic Egoism Behaviors in Large Language Models
von: Zhang, Yaoyuan, et al.
Veröffentlicht: (2025)