Alignment-Weighted DPO: A principled reasoning approach to improve safety alignment
Fuente:
arXiv
Guardado en:
| Autores principales: | Hu, Mengxuan, Datla, Vivek V., Kumar, Anoop, Guan, Zihan, Li, Sheng, Samuel, Alfy, Liu, Daben |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FB-RAG: Improving RAG with Forward and Backward Lookup
por: Chawla, Kushal, et al.
Publicado: (2025)
por: Chawla, Kushal, et al.
Publicado: (2025)
Confidence-Based Response Abstinence: Improving LLM Trustworthiness via Activation-Based Uncertainty Estimation
por: Huang, Zhiqi, et al.
Publicado: (2025)
por: Huang, Zhiqi, et al.
Publicado: (2025)
A Comparison of Independent and Joint Fine-tuning Strategies for Retrieval-Augmented Generation
por: Lawton, Neal Gregory, et al.
Publicado: (2025)
por: Lawton, Neal Gregory, et al.
Publicado: (2025)
Readability Reconsidered: A Cross-Dataset Analysis of Reference-Free Metrics
por: Belem, Catarina G, et al.
Publicado: (2025)
por: Belem, Catarina G, et al.
Publicado: (2025)
Play by the Type Rules: Inferring Constraints for LLM Functions in Declarative Programs
por: Glenn, Parker, et al.
Publicado: (2025)
por: Glenn, Parker, et al.
Publicado: (2025)
Improving Consistency in Retrieval-Augmented Systems with Group Similarity Rewards
por: Hamman, Faisal, et al.
Publicado: (2025)
por: Hamman, Faisal, et al.
Publicado: (2025)
LLM Optimization Unlocks Real-Time Pairwise Reranking
por: Wu, Jingyu, et al.
Publicado: (2025)
por: Wu, Jingyu, et al.
Publicado: (2025)
Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models
por: Jain, Neel, et al.
Publicado: (2024)
por: Jain, Neel, et al.
Publicado: (2024)
Are LLMs Ready for Neural-integrated Mechanistic Modeling? A Benchmark and Agentic Framework
por: Guan, Zihan, et al.
Publicado: (2026)
por: Guan, Zihan, et al.
Publicado: (2026)
Cat-DPO: Category-Adaptive Safety Alignment
por: Yang, Tiankai, et al.
Publicado: (2026)
por: Yang, Tiankai, et al.
Publicado: (2026)
Harmonizing Diverse Models: A Layer-wise Merging Strategy for Consistent Generation
por: Peng, Xujun, et al.
Publicado: (2025)
por: Peng, Xujun, et al.
Publicado: (2025)
MixDPO: Modeling Preference Strength for Pluralistic Alignment
por: Imai, Saki, et al.
Publicado: (2026)
por: Imai, Saki, et al.
Publicado: (2026)
Benign Samples Matter! Fine-tuning On Outlier Benign Samples Severely Breaks Safety
por: Guan, Zihan, et al.
Publicado: (2025)
por: Guan, Zihan, et al.
Publicado: (2025)
EPSVec: Efficient and Private Synthetic Data Generation via Dataset Vectors
por: Banayeeanzade, Amin, et al.
Publicado: (2026)
por: Banayeeanzade, Amin, et al.
Publicado: (2026)
A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
por: Lee, Andrew, et al.
Publicado: (2024)
por: Lee, Andrew, et al.
Publicado: (2024)
Uncertainty as Feature Gaps: Epistemic Uncertainty Quantification of LLMs in Contextual Question-Answering
por: Bakman, Yavuz, et al.
Publicado: (2025)
por: Bakman, Yavuz, et al.
Publicado: (2025)
Large Language Models for Causal Discovery: Current Landscape and Future Directions
por: Wan, Guangya, et al.
Publicado: (2024)
por: Wan, Guangya, et al.
Publicado: (2024)
Curry-DPO: Enhancing Alignment using Curriculum Learning & Ranked Preferences
por: Pattnaik, Pulkit, et al.
Publicado: (2024)
por: Pattnaik, Pulkit, et al.
Publicado: (2024)
BalancEdit: Dynamically Balancing the Generality-Locality Trade-off in Multi-modal Model Editing
por: Guo, Dongliang, et al.
Publicado: (2025)
por: Guo, Dongliang, et al.
Publicado: (2025)
SP^2DPO: An LLM-assisted Semantic Per-Pair DPO Generalization
por: He, Chaoyue, et al.
Publicado: (2026)
por: He, Chaoyue, et al.
Publicado: (2026)
Context-DPO: Aligning Language Models for Context-Faithfulness
por: Bi, Baolong, et al.
Publicado: (2024)
por: Bi, Baolong, et al.
Publicado: (2024)
Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment
por: Tice, Cameron, et al.
Publicado: (2026)
por: Tice, Cameron, et al.
Publicado: (2026)
daDPO: Distribution-Aware DPO for Distilling Conversational Abilities
por: Zhang, Zhengze, et al.
Publicado: (2025)
por: Zhang, Zhengze, et al.
Publicado: (2025)
TruthTorchLM: A Comprehensive Library for Predicting Truthfulness in LLM Outputs
por: Yaldiz, Duygu Nur, et al.
Publicado: (2025)
por: Yaldiz, Duygu Nur, et al.
Publicado: (2025)
Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive
por: Pal, Arka, et al.
Publicado: (2024)
por: Pal, Arka, et al.
Publicado: (2024)
Enhanced Diagnostic Performance via Large-Resolution Inference Optimization for Pathology Foundation Models
por: Hu, Mengxuan, et al.
Publicado: (2026)
por: Hu, Mengxuan, et al.
Publicado: (2026)
Temperature-scaling surprisal estimates improve fit to human reading times -- but does it do so for the "right reasons"?
por: Liu, Tong, et al.
Publicado: (2023)
por: Liu, Tong, et al.
Publicado: (2023)
Teaching with Lies: Curriculum DPO on Synthetic Negatives for Hallucination Detection
por: Pandit, Shrey, et al.
Publicado: (2025)
por: Pandit, Shrey, et al.
Publicado: (2025)
Aligning Large Language Models with Counterfactual DPO
por: Butcher, Bradley
Publicado: (2024)
por: Butcher, Bradley
Publicado: (2024)
Backdoor in Seconds: Unlocking Vulnerabilities in Large Pre-trained Models via Model Editing
por: Guo, Dongliang, et al.
Publicado: (2024)
por: Guo, Dongliang, et al.
Publicado: (2024)
Rethinking DPO: The Role of Rejected Responses in Preference Misalignment
por: Cho, Jay Hyeon, et al.
Publicado: (2025)
por: Cho, Jay Hyeon, et al.
Publicado: (2025)
No Free Lunch: Retrieval-Augmented Generation Undermines Fairness in LLMs, Even for Vigilant Users
por: Hu, Mengxuan, et al.
Publicado: (2024)
por: Hu, Mengxuan, et al.
Publicado: (2024)
Polypersona: Persona-Grounded LLM for Synthetic Survey Responses
por: Dash, Tejaswani, et al.
Publicado: (2025)
por: Dash, Tejaswani, et al.
Publicado: (2025)
Recovering Diversity Without Losing Alignment: A DPO Recipe for Post-Trained LLMs
por: Samuel, Vinay, et al.
Publicado: (2026)
por: Samuel, Vinay, et al.
Publicado: (2026)
Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective
por: Feng, Duanyu, et al.
Publicado: (2024)
por: Feng, Duanyu, et al.
Publicado: (2024)
sDPO: Don't Use Your Data All at Once
por: Kim, Dahyun, et al.
Publicado: (2024)
por: Kim, Dahyun, et al.
Publicado: (2024)
2D-DPO: Scaling Direct Preference Optimization with 2-Dimensional Supervision
por: Li, Shilong, et al.
Publicado: (2024)
por: Li, Shilong, et al.
Publicado: (2024)
BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning
por: Zhang, Beichen, et al.
Publicado: (2025)
por: Zhang, Beichen, et al.
Publicado: (2025)
Meaningless is better: hashing bias-inducing words in LLM prompts improves performance in logical reasoning and statistical learning
por: Chadimová, Milena, et al.
Publicado: (2024)
por: Chadimová, Milena, et al.
Publicado: (2024)
Multi-step retrieval and reasoning improves radiology question answering with large language models
por: Wind, Sebastian, et al.
Publicado: (2025)
por: Wind, Sebastian, et al.
Publicado: (2025)
Ejemplares similares
-
FB-RAG: Improving RAG with Forward and Backward Lookup
por: Chawla, Kushal, et al.
Publicado: (2025) -
Confidence-Based Response Abstinence: Improving LLM Trustworthiness via Activation-Based Uncertainty Estimation
por: Huang, Zhiqi, et al.
Publicado: (2025) -
A Comparison of Independent and Joint Fine-tuning Strategies for Retrieval-Augmented Generation
por: Lawton, Neal Gregory, et al.
Publicado: (2025) -
Readability Reconsidered: A Cross-Dataset Analysis of Reference-Free Metrics
por: Belem, Catarina G, et al.
Publicado: (2025) -
Play by the Type Rules: Inferring Constraints for LLM Functions in Declarative Programs
por: Glenn, Parker, et al.
Publicado: (2025)