A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Lee, Andrew, Bai, Xiaoyan, Pres, Itamar, Wattenberg, Martin, Kummerfeld, Jonathan K., Mihalcea, Rada |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Why Can't Transformers Learn Multiplication? Reverse-Engineering Reveals Long-Range Dependency Pitfalls
par: Bai, Xiaoyan, et autres
Publié: (2025)
par: Bai, Xiaoyan, et autres
Publié: (2025)
Towards Algorithmic Fidelity: Mental Health Representation across Demographics in Synthetic vs. Human-generated Data
par: Mori, Shinka, et autres
Publié: (2024)
par: Mori, Shinka, et autres
Publié: (2024)
Rethinking Table Instruction Tuning
par: Deng, Naihao, et autres
Publié: (2025)
par: Deng, Naihao, et autres
Publié: (2025)
Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
par: Pres, Itamar, et autres
Publié: (2024)
par: Pres, Itamar, et autres
Publié: (2024)
Chord Embeddings: Analyzing What They Capture and Their Role for Next Chord Prediction and Artist Attribute Prediction
par: Lahnala, Allison, et autres
Publié: (2021)
par: Lahnala, Allison, et autres
Publié: (2021)
Mind the (Belief) Gap: Group Identity in the World of LLMs
par: Borah, Angana, et autres
Publié: (2025)
par: Borah, Angana, et autres
Publié: (2025)
MAiDE-up: Multilingual Deception Detection of GPT-generated Hotel Reviews
par: Ignat, Oana, et autres
Publié: (2024)
par: Ignat, Oana, et autres
Publié: (2024)
Cross-cultural Inspiration Detection and Analysis in Real and LLM-generated Social Media Data
par: Ignat, Oana, et autres
Publié: (2024)
par: Ignat, Oana, et autres
Publié: (2024)
The Power of Many: Multi-Agent Multimodal Models for Cultural Image Captioning
par: Bai, Longju, et autres
Publié: (2024)
par: Bai, Longju, et autres
Publié: (2024)
Annotations on a Budget: Leveraging Geo-Data Similarity to Balance Model Performance and Annotation Cost
par: Ignat, Oana, et autres
Publié: (2024)
par: Ignat, Oana, et autres
Publié: (2024)
Chumor 2.0: Towards Benchmarking Chinese Humor Understanding
par: He, Ruiqi, et autres
Publié: (2024)
par: He, Ruiqi, et autres
Publié: (2024)
Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons
par: Chen, Jianhui, et autres
Publié: (2024)
par: Chen, Jianhui, et autres
Publié: (2024)
Belief-Sim: Towards Belief-Driven Simulation of Demographic Misinformation Susceptibility
par: Borah, Angana, et autres
Publié: (2026)
par: Borah, Angana, et autres
Publié: (2026)
Cat-DPO: Category-Adaptive Safety Alignment
par: Yang, Tiankai, et autres
Publié: (2026)
par: Yang, Tiankai, et autres
Publié: (2026)
VERVE: Template-based ReflectiVE Rewriting for MotiVational IntErviewing
par: Min, Do June, et autres
Publié: (2023)
par: Min, Do June, et autres
Publié: (2023)
Aligning AI Research with the Needs of Clinical Coding Workflows: Eight Recommendations Based on US Data Analysis and Critical Review
par: Gan, Yidong, et autres
Publié: (2024)
par: Gan, Yidong, et autres
Publié: (2024)
Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective
par: Feng, Duanyu, et autres
Publié: (2024)
par: Feng, Duanyu, et autres
Publié: (2024)
Alignment-Weighted DPO: A principled reasoning approach to improve safety alignment
par: Hu, Mengxuan, et autres
Publié: (2026)
par: Hu, Mengxuan, et autres
Publié: (2026)
Simple and Effective Baselines for Code Summarisation Evaluation
par: Robinson, Jade, et autres
Publié: (2025)
par: Robinson, Jade, et autres
Publié: (2025)
What Does it Mean for a Neural Network to Learn a "World Model"?
par: Li, Kenneth, et autres
Publié: (2025)
par: Li, Kenneth, et autres
Publié: (2025)
Eeyore: Realistic Depression Simulation via Supervised and Preference Optimization
par: Liu, Siyang, et autres
Publié: (2025)
par: Liu, Siyang, et autres
Publié: (2025)
Are Human Interactions Replicable by Generative Agents? A Case Study on Pronoun Usage in Hierarchical Interactions
par: Deng, Naihao, et autres
Publié: (2025)
par: Deng, Naihao, et autres
Publié: (2025)
Benchmarking and Improving LLM Robustness for Personalized Generation
par: Okite, Chimaobi, et autres
Publié: (2025)
par: Okite, Chimaobi, et autres
Publié: (2025)
Building Resource-Constrained Language Agents: A Korean Case Study on Chemical Toxicity Information
par: Cho, Hojun, et autres
Publié: (2025)
par: Cho, Hojun, et autres
Publié: (2025)
Adversarial DPO: Harnessing Harmful Data for Reducing Toxicity with Minimal Impact on Coherence and Evasiveness in Dialogue Agents
par: Kim, San, et autres
Publié: (2024)
par: Kim, San, et autres
Publié: (2024)
Culture Affordance Atlas: Reconciling Object Diversity Through Functional Mapping
par: Nwatu, Joan, et autres
Publié: (2025)
par: Nwatu, Joan, et autres
Publié: (2025)
MixDPO: Modeling Preference Strength for Pluralistic Alignment
par: Imai, Saki, et autres
Publié: (2026)
par: Imai, Saki, et autres
Publié: (2026)
$R^3$: "This is My SQL, Are You With Me?" A Consensus-Based Multi-Agent System for Text-to-SQL Tasks
par: Xia, Hanchen, et autres
Publié: (2024)
par: Xia, Hanchen, et autres
Publié: (2024)
The Generation Gap: Exploring Age Bias in the Value Systems of Large Language Models
par: Liu, Siyang, et autres
Publié: (2024)
par: Liu, Siyang, et autres
Publié: (2024)
Understanding or Memorizing? A Case Study of German Definite Articles in Language Models
par: Drechsel, Jonathan, et autres
Publié: (2026)
par: Drechsel, Jonathan, et autres
Publié: (2026)
Deception Detection from Linguistic and Physiological Data Streams Using Bimodal Convolutional Neural Networks
par: Li, Panfeng, et autres
Publié: (2023)
par: Li, Panfeng, et autres
Publié: (2023)
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
par: Pandey, Punya Syon, et autres
Publié: (2025)
par: Pandey, Punya Syon, et autres
Publié: (2025)
How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis
par: Yang, Yushi, et autres
Publié: (2024)
par: Yang, Yushi, et autres
Publié: (2024)
Curry-DPO: Enhancing Alignment using Curriculum Learning & Ranked Preferences
par: Pattnaik, Pulkit, et autres
Publié: (2024)
par: Pattnaik, Pulkit, et autres
Publié: (2024)
Rethinking DPO: The Role of Rejected Responses in Preference Misalignment
par: Cho, Jay Hyeon, et autres
Publié: (2025)
par: Cho, Jay Hyeon, et autres
Publié: (2025)
GloSS over Toxicity: Understanding and Mitigating Toxicity in LLMs via Global Toxic Subspace
par: Duan, Zenghao, et autres
Publié: (2025)
par: Duan, Zenghao, et autres
Publié: (2025)
Induction Head Toxicity Mechanistically Explains Repetition Curse in Large Language Models
par: Wang, Shuxun, et autres
Publié: (2025)
par: Wang, Shuxun, et autres
Publié: (2025)
When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas
par: Backmann, Steffen, et autres
Publié: (2025)
par: Backmann, Steffen, et autres
Publié: (2025)
An Empirical Study of SFT-DPO Interaction and Parameterization in Small Language Models
par: Feng, Yuming, et autres
Publié: (2026)
par: Feng, Yuming, et autres
Publié: (2026)
daDPO: Distribution-Aware DPO for Distilling Conversational Abilities
par: Zhang, Zhengze, et autres
Publié: (2025)
par: Zhang, Zhengze, et autres
Publié: (2025)
Documents similaires
-
Why Can't Transformers Learn Multiplication? Reverse-Engineering Reveals Long-Range Dependency Pitfalls
par: Bai, Xiaoyan, et autres
Publié: (2025) -
Towards Algorithmic Fidelity: Mental Health Representation across Demographics in Synthetic vs. Human-generated Data
par: Mori, Shinka, et autres
Publié: (2024) -
Rethinking Table Instruction Tuning
par: Deng, Naihao, et autres
Publié: (2025) -
Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
par: Pres, Itamar, et autres
Publié: (2024) -
Chord Embeddings: Analyzing What They Capture and Their Role for Next Chord Prediction and Artist Attribute Prediction
par: Lahnala, Allison, et autres
Publié: (2021)