Lightweight Robust Direct Preference Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Cheol Woo, Verma, Shresth, Tec, Mauricio, Tambe, Milind |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Preference Robustness for DPO with Applications to Public Health
by: Kim, Cheol Woo, et al.
Published: (2025)
by: Kim, Cheol Woo, et al.
Published: (2025)
Navigating the Social Welfare Frontier: Portfolios for Multi-objective Reinforcement Learning
by: Kim, Cheol Woo, et al.
Published: (2025)
by: Kim, Cheol Woo, et al.
Published: (2025)
Balancing Act: Prioritization Strategies for LLM-Designed Restless Bandit Rewards
by: Verma, Shresth, et al.
Published: (2024)
by: Verma, Shresth, et al.
Published: (2024)
Decisions and Deployment: The Five-Year SAHELI Project (2020-2025) on Restless Multi-Armed Bandits for Improving Maternal and Child Health
by: Verma, Shresth, et al.
Published: (2026)
by: Verma, Shresth, et al.
Published: (2026)
Rule-Bottleneck Reinforcement Learning: Joint Explanation and Decision Optimization for Resource Allocation with Language Agents
by: Tec, Mauricio, et al.
Published: (2025)
by: Tec, Mauricio, et al.
Published: (2025)
A Machine Learning Approach to Two-Stage Adaptive Robust Optimization
by: Bertsimas, Dimitris, et al.
Published: (2023)
by: Bertsimas, Dimitris, et al.
Published: (2023)
Analyzing Cost-Sensitive Surrogate Losses via $\mathcal{H}$-calibration
by: Shah, Sanket, et al.
Published: (2025)
by: Shah, Sanket, et al.
Published: (2025)
Leaving the Nest: Going Beyond Local Loss Functions for Predict-Then-Optimize
by: Shah, Sanket, et al.
Published: (2023)
by: Shah, Sanket, et al.
Published: (2023)
Improving Health Information Access in the World's Largest Maternal Mobile Health Program via Bandit Algorithms
by: Lalan, Arshika, et al.
Published: (2024)
by: Lalan, Arshika, et al.
Published: (2024)
Combining Diverse Information for Coordinated Action: Stochastic Bandit Algorithms for Heterogeneous Agents
by: Gordon, Lucia, et al.
Published: (2024)
by: Gordon, Lucia, et al.
Published: (2024)
Efficient Public Health Intervention Planning Using Decomposition-Based Decision-Focused Learning
by: Shah, Sanket, et al.
Published: (2024)
by: Shah, Sanket, et al.
Published: (2024)
What is the Right Notion of Distance between Predict-then-Optimize Tasks?
by: Rodriguez-Diaz, Paula, et al.
Published: (2024)
by: Rodriguez-Diaz, Paula, et al.
Published: (2024)
Optimal Control of Multiclass Fluid Queueing Networks: A Machine Learning Approach
by: Bertsimas, Dimitris, et al.
Published: (2023)
by: Bertsimas, Dimitris, et al.
Published: (2023)
Reinforcement learning with combinatorial actions for coupled restless bandits
by: Xu, Lily, et al.
Published: (2025)
by: Xu, Lily, et al.
Published: (2025)
Context in Public Health for Underserved Communities: A Bayesian Approach to Online Restless Bandits
by: Liang, Biyonka, et al.
Published: (2024)
by: Liang, Biyonka, et al.
Published: (2024)
Robust LLM Alignment via Distributionally Robust Direct Preference Optimization
by: Xu, Zaiyan, et al.
Published: (2025)
by: Xu, Zaiyan, et al.
Published: (2025)
On Diffusion Models for Multi-Agent Partial Observability: Shared Attractors, Error Bounds, and Composite Flow
by: Wang, Tonghan, et al.
Published: (2024)
by: Wang, Tonghan, et al.
Published: (2024)
Understanding the Impact of Sampling Quality in Direct Preference Optimization
by: Kim, Kyung Rok, et al.
Published: (2025)
by: Kim, Kyung Rok, et al.
Published: (2025)
Many Preferences, Few Policies: Towards Scalable Language Model Personalization
by: Kim, Cheol Woo, et al.
Published: (2026)
by: Kim, Cheol Woo, et al.
Published: (2026)
Generative AI for Social Impact
by: Kong, Lingkai, et al.
Published: (2026)
by: Kong, Lingkai, et al.
Published: (2026)
Composite Flow Matching for Reinforcement Learning with Shifted-Dynamics Data
by: Kong, Lingkai, et al.
Published: (2025)
by: Kong, Lingkai, et al.
Published: (2025)
The Bandit Whisperer: Communication Learning for Restless Bandits
by: Zhao, Yunfan, et al.
Published: (2024)
by: Zhao, Yunfan, et al.
Published: (2024)
Improving the Prediction of Individual Engagement in Recommendations Using Cognitive Models
by: Seow, Roderick, et al.
Published: (2024)
by: Seow, Roderick, et al.
Published: (2024)
Contrasting local and global modeling with machine learning and satellite data: A case study estimating tree canopy height in African savannas
by: Rolf, Esther, et al.
Published: (2024)
by: Rolf, Esther, et al.
Published: (2024)
Distributed Direct Preference Optimization
by: Jiang, Zhanhong
Published: (2026)
by: Jiang, Zhanhong
Published: (2026)
CompassDPO: Dynamics-Controlled Direct Preference Optimization for Robust Safety Alignment
by: Liu, Jilong, et al.
Published: (2026)
by: Liu, Jilong, et al.
Published: (2026)
Artificial Replay: A Meta-Algorithm for Harnessing Historical Data in Bandits
by: Banerjee, Siddhartha, et al.
Published: (2022)
by: Banerjee, Siddhartha, et al.
Published: (2022)
Dual-Mandate Patrols: Multi-Armed Bandits for Green Security
by: Xu, Lily, et al.
Published: (2020)
by: Xu, Lily, et al.
Published: (2020)
Bayesian Collaborative Bandits with Thompson Sampling for Improved Outreach in Maternal Health Program
by: Dasgupta, Arpan, et al.
Published: (2024)
by: Dasgupta, Arpan, et al.
Published: (2024)
ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment
by: Lin, Xiaoqiang, et al.
Published: (2025)
by: Lin, Xiaoqiang, et al.
Published: (2025)
Optimizing Heat Alert Issuance with Reinforcement Learning
by: Considine, Ellen M., et al.
Published: (2023)
by: Considine, Ellen M., et al.
Published: (2023)
Causal Estimation of Exposure Shifts with Neural Networks
by: Tec, Mauricio, et al.
Published: (2023)
by: Tec, Mauricio, et al.
Published: (2023)
Direct Preference Optimization With Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences
by: Chidambaram, Keertana, et al.
Published: (2024)
by: Chidambaram, Keertana, et al.
Published: (2024)
Gradient Imbalance in Direct Preference Optimization
by: Ma, Qinwei, et al.
Published: (2025)
by: Ma, Qinwei, et al.
Published: (2025)
Active Learning for Direct Preference Optimization
by: Kveton, Branislav, et al.
Published: (2025)
by: Kveton, Branislav, et al.
Published: (2025)
A Survey of Direct Preference Optimization
by: Liu, Shunyu, et al.
Published: (2025)
by: Liu, Shunyu, et al.
Published: (2025)
Optimal Control of Fluid Restless Multi-armed Bandits: A Machine Learning Approach
by: Bertsimas, Dimitris, et al.
Published: (2025)
by: Bertsimas, Dimitris, et al.
Published: (2025)
Incentive-Aware AI Safety via Strategic Resource Allocation: A Stackelberg Security Games Perspective
by: Kim, Cheol Woo, et al.
Published: (2026)
by: Kim, Cheol Woo, et al.
Published: (2026)
IRL for Restless Multi-Armed Bandits with Applications in Maternal and Child Health
by: Jain, Gauri, et al.
Published: (2024)
by: Jain, Gauri, et al.
Published: (2024)
Reward Shaping for Inference-Time Alignment: A Stackelberg Game Perspective
by: Wang, Haichuan, et al.
Published: (2026)
by: Wang, Haichuan, et al.
Published: (2026)
Similar Items
-
Preference Robustness for DPO with Applications to Public Health
by: Kim, Cheol Woo, et al.
Published: (2025) -
Navigating the Social Welfare Frontier: Portfolios for Multi-objective Reinforcement Learning
by: Kim, Cheol Woo, et al.
Published: (2025) -
Balancing Act: Prioritization Strategies for LLM-Designed Restless Bandit Rewards
by: Verma, Shresth, et al.
Published: (2024) -
Decisions and Deployment: The Five-Year SAHELI Project (2020-2025) on Restless Multi-Armed Bandits for Improving Maternal and Child Health
by: Verma, Shresth, et al.
Published: (2026) -
Rule-Bottleneck Reinforcement Learning: Joint Explanation and Decision Optimization for Resource Allocation with Language Agents
by: Tec, Mauricio, et al.
Published: (2025)