Model Editing as a Robust and Denoised variant of DPO: A Case Study on Toxicity
Fuente:
arXiv
Guardado en:
| Autores principales: | Uppaal, Rheeya, Dey, Apratim, He, Yiting, Zhong, Yiqiao, Hu, Junjie |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
How Useful is Continued Pre-Training for Generative Unsupervised Domain Adaptation?
por: Uppaal, Rheeya, et al.
Publicado: (2024)
por: Uppaal, Rheeya, et al.
Publicado: (2024)
When Safety Fails Before the Answer: Benchmarking Harmful Behavior Detection in Reasoning Chains
por: Kakkar, Ishita, et al.
Publicado: (2026)
por: Kakkar, Ishita, et al.
Publicado: (2026)
Shattered Compositionality: Counterintuitive Learning Dynamics of Transformers for Arithmetic
por: Zhao, Xingyu, et al.
Publicado: (2026)
por: Zhao, Xingyu, et al.
Publicado: (2026)
Journey Before Destination: On the importance of Visual Faithfulness in Slow Thinking
por: Uppaal, Rheeya, et al.
Publicado: (2025)
por: Uppaal, Rheeya, et al.
Publicado: (2025)
A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
por: Lee, Andrew, et al.
Publicado: (2024)
por: Lee, Andrew, et al.
Publicado: (2024)
How does Multi-Task Training Affect Transformer In-Context Capabilities? Investigations with Function Classes
por: Bhasin, Harmon, et al.
Publicado: (2024)
por: Bhasin, Harmon, et al.
Publicado: (2024)
Improving Bilingual Capabilities of Language Models to Support Diverse Linguistic Practices in Education
por: Syamkumar, Anand, et al.
Publicado: (2024)
por: Syamkumar, Anand, et al.
Publicado: (2024)
How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis
por: Yang, Yushi, et al.
Publicado: (2024)
por: Yang, Yushi, et al.
Publicado: (2024)
Provably Robust DPO: Aligning Language Models with Noisy Feedback
por: Chowdhury, Sayak Ray, et al.
Publicado: (2024)
por: Chowdhury, Sayak Ray, et al.
Publicado: (2024)
V-SEAM: Visual Semantic Editing and Attention Modulating for Causal Interpretability of Vision-Language Models
por: Wang, Qidong, et al.
Publicado: (2025)
por: Wang, Qidong, et al.
Publicado: (2025)
Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study
por: Xu, Shusheng, et al.
Publicado: (2024)
por: Xu, Shusheng, et al.
Publicado: (2024)
Robust Multi-Objective Preference Alignment with Online DPO
por: Gupta, Raghav, et al.
Publicado: (2025)
por: Gupta, Raghav, et al.
Publicado: (2025)
On the Robustness of Knowledge Editing for Detoxification
por: Dong, Ming, et al.
Publicado: (2026)
por: Dong, Ming, et al.
Publicado: (2026)
SP^2DPO: An LLM-assisted Semantic Per-Pair DPO Generalization
por: He, Chaoyue, et al.
Publicado: (2026)
por: He, Chaoyue, et al.
Publicado: (2026)
An Empirical Study of SFT-DPO Interaction and Parameterization in Small Language Models
por: Feng, Yuming, et al.
Publicado: (2026)
por: Feng, Yuming, et al.
Publicado: (2026)
Task Vector Geometry Underlies Dual Modes of Task Inference in Transformers
por: Yan, Hao, et al.
Publicado: (2026)
por: Yan, Hao, et al.
Publicado: (2026)
On the Robustness of Editing Large Language Models
por: Ma, Xinbei, et al.
Publicado: (2024)
por: Ma, Xinbei, et al.
Publicado: (2024)
Green AI: Exploring Carbon Footprints, Mitigation Strategies, and Trade Offs in Large Language Model Training
por: Liu, Vivian, et al.
Publicado: (2024)
por: Liu, Vivian, et al.
Publicado: (2024)
DPO Meets PPO: Reinforced Token Optimization for RLHF
por: Zhong, Han, et al.
Publicado: (2024)
por: Zhong, Han, et al.
Publicado: (2024)
Out-of-distribution generalization via composition: a lens through induction heads in Transformers
por: Song, Jiajun, et al.
Publicado: (2024)
por: Song, Jiajun, et al.
Publicado: (2024)
Unifying Attention Heads and Task Vectors via Hidden State Geometry in In-Context Learning
por: Yang, Haolin, et al.
Publicado: (2025)
por: Yang, Haolin, et al.
Publicado: (2025)
An Information-Theoretic Framework for Robust Large Language Model Editing
por: Chen, Qizhou, et al.
Publicado: (2025)
por: Chen, Qizhou, et al.
Publicado: (2025)
Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation
por: Tu, Songjun, et al.
Publicado: (2025)
por: Tu, Songjun, et al.
Publicado: (2025)
Aligning Large Language Models with Counterfactual DPO
por: Butcher, Bradley
Publicado: (2024)
por: Butcher, Bradley
Publicado: (2024)
Bootstrapping Language Models with DPO Implicit Rewards
por: Chen, Changyu, et al.
Publicado: (2024)
por: Chen, Changyu, et al.
Publicado: (2024)
Analyzing Toxicity in Deep Conversations: A Reddit Case Study
por: Shankaran, Vigneshwaran, et al.
Publicado: (2024)
por: Shankaran, Vigneshwaran, et al.
Publicado: (2024)
Adversarial DPO: Harnessing Harmful Data for Reducing Toxicity with Minimal Impact on Coherence and Evasiveness in Dialogue Agents
por: Kim, San, et al.
Publicado: (2024)
por: Kim, San, et al.
Publicado: (2024)
daDPO: Distribution-Aware DPO for Distilling Conversational Abilities
por: Zhang, Zhengze, et al.
Publicado: (2025)
por: Zhang, Zhengze, et al.
Publicado: (2025)
D2PO: Discriminator-Guided DPO with Response Evaluation Models
por: Singhal, Prasann, et al.
Publicado: (2024)
por: Singhal, Prasann, et al.
Publicado: (2024)
DPO-Tuned Large Language Models for Segmentation in Simultaneous Speech Translation
por: Yang, Zeyu, et al.
Publicado: (2025)
por: Yang, Zeyu, et al.
Publicado: (2025)
Enhancing Text Editing for Grammatical Error Correction: Arabic as a Case Study
por: Alhafni, Bashar, et al.
Publicado: (2025)
por: Alhafni, Bashar, et al.
Publicado: (2025)
BiasDPO: Mitigating Bias in Language Models through Direct Preference Optimization
por: Allam, Ahmed
Publicado: (2024)
por: Allam, Ahmed
Publicado: (2024)
Adaptive Detoxification: Safeguarding General Capabilities of LLMs through Toxicity-Aware Knowledge Editing
por: Lu, Yifan, et al.
Publicado: (2025)
por: Lu, Yifan, et al.
Publicado: (2025)
DPO-Shift: Shifting the Distribution of Direct Preference Optimization
por: Yang, Xiliang, et al.
Publicado: (2025)
por: Yang, Xiliang, et al.
Publicado: (2025)
Alignment-Weighted DPO: A principled reasoning approach to improve safety alignment
por: Hu, Mengxuan, et al.
Publicado: (2026)
por: Hu, Mengxuan, et al.
Publicado: (2026)
Robust and Scalable Model Editing for Large Language Models
por: Chen, Yingfa, et al.
Publicado: (2024)
por: Chen, Yingfa, et al.
Publicado: (2024)
Context-Robust Knowledge Editing for Language Models
por: Park, Haewon, et al.
Publicado: (2025)
por: Park, Haewon, et al.
Publicado: (2025)
Pre-DPO: Improving Data Utilization in Direct Preference Optimization Using a Guiding Reference Model
por: Pan, Junshu, et al.
Publicado: (2025)
por: Pan, Junshu, et al.
Publicado: (2025)
EAMET: Robust Massive Model Editing via Embedding Alignment Optimization
por: Dai, Yanbo, et al.
Publicado: (2025)
por: Dai, Yanbo, et al.
Publicado: (2025)
Toxicity Detection for Free
por: Hu, Zhanhao, et al.
Publicado: (2024)
por: Hu, Zhanhao, et al.
Publicado: (2024)
Ejemplares similares
-
How Useful is Continued Pre-Training for Generative Unsupervised Domain Adaptation?
por: Uppaal, Rheeya, et al.
Publicado: (2024) -
When Safety Fails Before the Answer: Benchmarking Harmful Behavior Detection in Reasoning Chains
por: Kakkar, Ishita, et al.
Publicado: (2026) -
Shattered Compositionality: Counterintuitive Learning Dynamics of Transformers for Arithmetic
por: Zhao, Xingyu, et al.
Publicado: (2026) -
Journey Before Destination: On the importance of Visual Faithfulness in Slow Thinking
por: Uppaal, Rheeya, et al.
Publicado: (2025) -
A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
por: Lee, Andrew, et al.
Publicado: (2024)