Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?
Fuente:
arXiv
Guardado en:
| Autores principales: | Gölz, Paul, Haghtalab, Nika, Yang, Kunhe |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Is Knowledge Power? On the (Im)possibility of Learning from Strategic Interactions
por: Ananthakrishnan, Nivasini, et al.
Publicado: (2024)
por: Ananthakrishnan, Nivasini, et al.
Publicado: (2024)
Leakage-Robust Bayesian Persuasion
por: Haghtalab, Nika, et al.
Publicado: (2024)
por: Haghtalab, Nika, et al.
Publicado: (2024)
Platforms for Efficient and Incentive-Aware Collaboration
por: Haghtalab, Nika, et al.
Publicado: (2024)
por: Haghtalab, Nika, et al.
Publicado: (2024)
Pluralistic Leaderboards
por: Haghtalab, Nika, et al.
Publicado: (2026)
por: Haghtalab, Nika, et al.
Publicado: (2026)
Improved Bayes Risk Can Yield Reduced Social Welfare Under Competition
por: Jagadeesan, Meena, et al.
Publicado: (2023)
por: Jagadeesan, Meena, et al.
Publicado: (2023)
Smooth Nash Equilibria: Algorithms and Complexity
por: Daskalakis, Constantinos, et al.
Publicado: (2023)
por: Daskalakis, Constantinos, et al.
Publicado: (2023)
Learning in Stackelberg Games with Non-myopic Agents
por: Haghtalab, Nika, et al.
Publicado: (2022)
por: Haghtalab, Nika, et al.
Publicado: (2022)
Should Decision-Makers Reveal Classifiers in Online Strategic Classification?
por: Shao, Han, et al.
Publicado: (2025)
por: Shao, Han, et al.
Publicado: (2025)
Learning Local Stackelberg Equilibria from Repeated Interactions with a Learning Agent
por: Ananthakrishnan, Nivasini, et al.
Publicado: (2025)
por: Ananthakrishnan, Nivasini, et al.
Publicado: (2025)
Fundamental Bounds on Online Strategic Classification
por: Ahmadi, Saba, et al.
Publicado: (2023)
por: Ahmadi, Saba, et al.
Publicado: (2023)
Strategic Littlestone Dimension: Improved Bounds on Online Strategic Classification
por: Ahmadi, Saba, et al.
Publicado: (2024)
por: Ahmadi, Saba, et al.
Publicado: (2024)
Fundamental Limits of Game-Theoretic LLM Alignment: Smith Consistency and Preference Matching
por: Shi, Zhekun, et al.
Publicado: (2025)
por: Shi, Zhekun, et al.
Publicado: (2025)
Response Time Enhances Alignment with Heterogeneous Preferences
por: Echenique, Federico, et al.
Publicado: (2026)
por: Echenique, Federico, et al.
Publicado: (2026)
Metric Distortion with Preference Intensities
por: Abbaszadeh, Mehrad, et al.
Publicado: (2026)
por: Abbaszadeh, Mehrad, et al.
Publicado: (2026)
Accelerated Preference Elicitation with LLM-Based Proxies
por: Huang, David, et al.
Publicado: (2025)
por: Huang, David, et al.
Publicado: (2025)
Regularized Online RLHF with Generalized Bilinear Preferences
por: Lee, Junghyun, et al.
Publicado: (2026)
por: Lee, Junghyun, et al.
Publicado: (2026)
Proportional Aggregation of Preferences for Sequential Decision Making
por: Chandak, Nikhil, et al.
Publicado: (2023)
por: Chandak, Nikhil, et al.
Publicado: (2023)
Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
por: Zhang, Yuheng, et al.
Publicado: (2024)
por: Zhang, Yuheng, et al.
Publicado: (2024)
Corruption-Robust Offline Two-Player Zero-Sum Markov Games
por: Nika, Andi, et al.
Publicado: (2024)
por: Nika, Andi, et al.
Publicado: (2024)
Stackelberg Learning from Human Feedback: Preference Optimization as a Sequential Game
por: Pásztor, Barna, et al.
Publicado: (2025)
por: Pásztor, Barna, et al.
Publicado: (2025)
Computational Aspects of Bayesian Persuasion under Approximate Best Response
por: Yang, Kunhe, et al.
Publicado: (2024)
por: Yang, Kunhe, et al.
Publicado: (2024)
LLM-Powered Preference Elicitation in Combinatorial Assignment
por: Soumalias, Ermis, et al.
Publicado: (2025)
por: Soumalias, Ermis, et al.
Publicado: (2025)
Online Recommendations for Agents with Discounted Adaptive Preferences
por: Agarwal, Arpit, et al.
Publicado: (2023)
por: Agarwal, Arpit, et al.
Publicado: (2023)
Finding Common Ground in a Sea of Alternatives
por: Chooi, Jay, et al.
Publicado: (2026)
por: Chooi, Jay, et al.
Publicado: (2026)
Is Online Linear Optimization Sufficient for Strategic Robustness?
por: Cai, Yang, et al.
Publicado: (2026)
por: Cai, Yang, et al.
Publicado: (2026)
Bandits with Preference Feedback: A Stackelberg Game Perspective
por: Pásztor, Barna, et al.
Publicado: (2024)
por: Pásztor, Barna, et al.
Publicado: (2024)
On the Distortion of Committee Election with 1-Euclidean Preferences and Few Distance Queries
por: Fotakis, Dimitris, et al.
Publicado: (2024)
por: Fotakis, Dimitris, et al.
Publicado: (2024)
Strategic Filtering for Content Moderation: Free Speech or Free of Distortion?
por: Ahmadi, Saba, et al.
Publicado: (2025)
por: Ahmadi, Saba, et al.
Publicado: (2025)
The Limits of Preference Data for Post-Training
por: Zhao, Eric, et al.
Publicado: (2025)
por: Zhao, Eric, et al.
Publicado: (2025)
How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators
por: Liu, Shang, et al.
Publicado: (2025)
por: Liu, Shang, et al.
Publicado: (2025)
User Response in Ad Auctions: An MDP Formulation of Long-Term Revenue Optimization
por: Cai, Yang, et al.
Publicado: (2023)
por: Cai, Yang, et al.
Publicado: (2023)
Communicating with Anecdotes
por: Haghtalab, Nika, et al.
Publicado: (2022)
por: Haghtalab, Nika, et al.
Publicado: (2022)
Two-sided Competing Matching Recommendation Markets With Quota and Complementary Preferences Constraints
por: Li, Yuantong, et al.
Publicado: (2023)
por: Li, Yuantong, et al.
Publicado: (2023)
COMAL: A Convergent Meta-Algorithm for Aligning LLMs with General Preferences
por: Liu, Yixin, et al.
Publicado: (2024)
por: Liu, Yixin, et al.
Publicado: (2024)
Online Stackelberg Optimization via Nonlinear Control
por: Brown, William, et al.
Publicado: (2024)
por: Brown, William, et al.
Publicado: (2024)
Maximally Random Sortition
por: de Azevedo, Gabriel, et al.
Publicado: (2026)
por: de Azevedo, Gabriel, et al.
Publicado: (2026)
Incentivizing Honesty among Competitors in Collaborative Learning and Optimization
por: Dorner, Florian E., et al.
Publicado: (2023)
por: Dorner, Florian E., et al.
Publicado: (2023)
Preferences Evolve And So Should Your Bandits: Bandits with Evolving States for Online Platforms
por: Khosravi, Khashayar, et al.
Publicado: (2023)
por: Khosravi, Khashayar, et al.
Publicado: (2023)
Generative Social Choice
por: Fish, Sara, et al.
Publicado: (2023)
por: Fish, Sara, et al.
Publicado: (2023)
Nash CoT: Multi-Path Inference with Preference Equilibrium
por: Zhang, Ziqi, et al.
Publicado: (2024)
por: Zhang, Ziqi, et al.
Publicado: (2024)
Ejemplares similares
-
Is Knowledge Power? On the (Im)possibility of Learning from Strategic Interactions
por: Ananthakrishnan, Nivasini, et al.
Publicado: (2024) -
Leakage-Robust Bayesian Persuasion
por: Haghtalab, Nika, et al.
Publicado: (2024) -
Platforms for Efficient and Incentive-Aware Collaboration
por: Haghtalab, Nika, et al.
Publicado: (2024) -
Pluralistic Leaderboards
por: Haghtalab, Nika, et al.
Publicado: (2026) -
Improved Bayes Risk Can Yield Reduced Social Welfare Under Competition
por: Jagadeesan, Meena, et al.
Publicado: (2023)