MixDPO: Modeling Preference Strength for Pluralistic Alignment
Fuente:
arXiv
Guardado en:
| Autores principales: | Imai, Saki, Heydari, Pedram, Sicilia, Anthony, Kaeberlein, Asteria, Atwell, Katherine, Alikhani, Malihe |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
BASIL: Bayesian Assessment of Sycophancy in LLMs
por: Atwell, Katherine, et al.
Publicado: (2025)
por: Atwell, Katherine, et al.
Publicado: (2025)
SiLVERScore: Semantically-Aware Embeddings for Sign Language Generation Evaluation
por: Imai, Saki, et al.
Publicado: (2025)
por: Imai, Saki, et al.
Publicado: (2025)
Measuring How (Not Just Whether) VLMs Build Common Ground
por: Imai, Saki, et al.
Publicado: (2025)
por: Imai, Saki, et al.
Publicado: (2025)
Evaluating Theory of (an uncertain) Mind: Predicting the Uncertain Beliefs of Others in Conversation Forecasting
por: Sicilia, Anthony, et al.
Publicado: (2024)
por: Sicilia, Anthony, et al.
Publicado: (2024)
Eliciting Uncertainty in Chain-of-Thought to Mitigate Bias against Forecasting Harmful User Behaviors
por: Sicilia, Anthony, et al.
Publicado: (2024)
por: Sicilia, Anthony, et al.
Publicado: (2024)
Deal, or no deal (or who knows)? Forecasting Uncertainty in Conversations using Large Language Models
por: Sicilia, Anthony, et al.
Publicado: (2024)
por: Sicilia, Anthony, et al.
Publicado: (2024)
Generating Signed Language Instructions in Large-Scale Dialogue Systems
por: İnan, Mert, et al.
Publicado: (2024)
por: İnan, Mert, et al.
Publicado: (2024)
Contextual ASR Error Handling with LLMs Augmentation for Goal-Oriented Conversational AI
por: Asano, Yuya, et al.
Publicado: (2025)
por: Asano, Yuya, et al.
Publicado: (2025)
Accounting for Sycophancy in Language Model Uncertainty Estimation
por: Sicilia, Anthony, et al.
Publicado: (2024)
por: Sicilia, Anthony, et al.
Publicado: (2024)
HumBEL: A Human-in-the-Loop Approach for Evaluating Demographic Factors of Language Models in Human-Machine Conversations
por: Sicilia, Anthony, et al.
Publicado: (2023)
por: Sicilia, Anthony, et al.
Publicado: (2023)
Curry-DPO: Enhancing Alignment using Curriculum Learning & Ranked Preferences
por: Pattnaik, Pulkit, et al.
Publicado: (2024)
por: Pattnaik, Pulkit, et al.
Publicado: (2024)
Mix- and MoE-DPO: A Variational Inference Approach to Direct Preference Optimization
por: Bohne, Jason, et al.
Publicado: (2025)
por: Bohne, Jason, et al.
Publicado: (2025)
Pluralistic Alignment for Healthcare: A Role-Driven Framework
por: Zhong, Jiayou, et al.
Publicado: (2025)
por: Zhong, Jiayou, et al.
Publicado: (2025)
"Nothing about us without us": Perspectives of Global Deaf and Hard-of-hearing Community Members on Sign Language Technologies
por: Atwell, Katherine, et al.
Publicado: (2025)
por: Atwell, Katherine, et al.
Publicado: (2025)
VITAL: A New Dataset for Benchmarking Pluralistic Alignment in Healthcare
por: Shetty, Anudeex, et al.
Publicado: (2025)
por: Shetty, Anudeex, et al.
Publicado: (2025)
VISPA: Pluralistic Alignment via Automatic Value Selection and Activation
por: Zheng, Shenyan, et al.
Publicado: (2026)
por: Zheng, Shenyan, et al.
Publicado: (2026)
AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization
por: Wu, Junkang, et al.
Publicado: (2024)
por: Wu, Junkang, et al.
Publicado: (2024)
Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive
por: Pal, Arka, et al.
Publicado: (2024)
por: Pal, Arka, et al.
Publicado: (2024)
A Persona-Based Evaluation Framework for Pluralistic Alignment in Generative AI
por: Karagoz, Atahan
Publicado: (2026)
por: Karagoz, Atahan
Publicado: (2026)
Studying and Mitigating Biases in Sign Language Understanding Models
por: Atwell, Katherine, et al.
Publicado: (2024)
por: Atwell, Katherine, et al.
Publicado: (2024)
Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap
por: Qi, Xuan, et al.
Publicado: (2025)
por: Qi, Xuan, et al.
Publicado: (2025)
Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing
por: Qi, Biqing, et al.
Publicado: (2024)
por: Qi, Biqing, et al.
Publicado: (2024)
Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering
por: Mohamed, Anas, et al.
Publicado: (2025)
por: Mohamed, Anas, et al.
Publicado: (2025)
Multi-Preference Optimization: Generalizing DPO via Set-Level Contrasts
por: Gupta, Taneesh, et al.
Publicado: (2024)
por: Gupta, Taneesh, et al.
Publicado: (2024)
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
por: Lai, Xin, et al.
Publicado: (2024)
por: Lai, Xin, et al.
Publicado: (2024)
APPA: Adaptive Preference Pluralistic Alignment for Fair Federated RLHF of LLMs
por: Srewa, Mahmoud, et al.
Publicado: (2026)
por: Srewa, Mahmoud, et al.
Publicado: (2026)
daDPO: Distribution-Aware DPO for Distilling Conversational Abilities
por: Zhang, Zhengze, et al.
Publicado: (2025)
por: Zhang, Zhengze, et al.
Publicado: (2025)
RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models
por: Khaki, Saeed, et al.
Publicado: (2024)
por: Khaki, Saeed, et al.
Publicado: (2024)
ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment
por: Lin, Xiaoqiang, et al.
Publicado: (2025)
por: Lin, Xiaoqiang, et al.
Publicado: (2025)
SP^2DPO: An LLM-assisted Semantic Per-Pair DPO Generalization
por: He, Chaoyue, et al.
Publicado: (2026)
por: He, Chaoyue, et al.
Publicado: (2026)
An Active Learning Framework for Inclusive Generation by Large Language Models
por: Hassan, Sabit, et al.
Publicado: (2024)
por: Hassan, Sabit, et al.
Publicado: (2024)
mDPO: Conditional Preference Optimization for Multimodal Large Language Models
por: Wang, Fei, et al.
Publicado: (2024)
por: Wang, Fei, et al.
Publicado: (2024)
Accelerated Preference Optimization for Large Language Model Alignment
por: He, Jiafan, et al.
Publicado: (2024)
por: He, Jiafan, et al.
Publicado: (2024)
Self-Play Preference Optimization for Language Model Alignment
por: Wu, Yue, et al.
Publicado: (2024)
por: Wu, Yue, et al.
Publicado: (2024)
Adaptive Alignment: Dynamic Preference Adjustments via Multi-Objective Reinforcement Learning for Pluralistic AI
por: Harland, Hadassah, et al.
Publicado: (2024)
por: Harland, Hadassah, et al.
Publicado: (2024)
Pairwise Calibrated Rewards for Pluralistic Alignment
por: Halpern, Daniel, et al.
Publicado: (2025)
por: Halpern, Daniel, et al.
Publicado: (2025)
When Models Examine Themselves: Vocabulary-Activation Correspondence in Self-Referential Processing
por: Dadfar, Zachary Pedram
Publicado: (2026)
por: Dadfar, Zachary Pedram
Publicado: (2026)
Hakim: Farsi Text Embedding Model
por: Sarmadi, Mehran, et al.
Publicado: (2025)
por: Sarmadi, Mehran, et al.
Publicado: (2025)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
por: Kim, Dongyoung, et al.
Publicado: (2024)
por: Kim, Dongyoung, et al.
Publicado: (2024)
Group Preference Optimization: Few-Shot Alignment of Large Language Models
por: Zhao, Siyan, et al.
Publicado: (2023)
por: Zhao, Siyan, et al.
Publicado: (2023)
Ejemplares similares
-
BASIL: Bayesian Assessment of Sycophancy in LLMs
por: Atwell, Katherine, et al.
Publicado: (2025) -
SiLVERScore: Semantically-Aware Embeddings for Sign Language Generation Evaluation
por: Imai, Saki, et al.
Publicado: (2025) -
Measuring How (Not Just Whether) VLMs Build Common Ground
por: Imai, Saki, et al.
Publicado: (2025) -
Evaluating Theory of (an uncertain) Mind: Predicting the Uncertain Beliefs of Others in Conversation Forecasting
por: Sicilia, Anthony, et al.
Publicado: (2024) -
Eliciting Uncertainty in Chain-of-Thought to Mitigate Bias against Forecasting Harmful User Behaviors
por: Sicilia, Anthony, et al.
Publicado: (2024)