How Sampling Shapes LLM Alignment: From One-Shot Optima to Iterative Dynamics
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Yurong, He, Yu, Jordan, Michael I., Yao, Fan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Calibeating Made Simple
di: Chen, Yurong, et al.
Pubblicazione: (2026)
di: Chen, Yurong, et al.
Pubblicazione: (2026)
Response Time Enhances Alignment with Heterogeneous Preferences
di: Echenique, Federico, et al.
Pubblicazione: (2026)
di: Echenique, Federico, et al.
Pubblicazione: (2026)
One-Shot Strategic Classification Under Unknown Costs
di: Rosenfeld, Elan, et al.
Pubblicazione: (2023)
di: Rosenfeld, Elan, et al.
Pubblicazione: (2023)
From Average-Iterate to Last-Iterate Convergence in Games: A Reduction and Its Applications
di: Cai, Yang, et al.
Pubblicazione: (2025)
di: Cai, Yang, et al.
Pubblicazione: (2025)
Last-Iterate Convergence of Adaptive Riemannian Gradient Descent for Equilibrium Computation
di: Cai, Yang, et al.
Pubblicazione: (2023)
di: Cai, Yang, et al.
Pubblicazione: (2023)
Fair Allocation in Dynamic Mechanism Design
di: Fallah, Alireza, et al.
Pubblicazione: (2024)
di: Fallah, Alireza, et al.
Pubblicazione: (2024)
Defection-Free Collaboration between Competitors in a Learning System
di: Werner, Mariel, et al.
Pubblicazione: (2024)
di: Werner, Mariel, et al.
Pubblicazione: (2024)
Learning Zero-Sum Linear Quadratic Games with Improved Sample Complexity and Last-Iterate Convergence
di: Wu, Jiduan, et al.
Pubblicazione: (2023)
di: Wu, Jiduan, et al.
Pubblicazione: (2023)
Fundamental Limits of Game-Theoretic LLM Alignment: Smith Consistency and Preference Matching
di: Shi, Zhekun, et al.
Pubblicazione: (2025)
di: Shi, Zhekun, et al.
Pubblicazione: (2025)
A Reinforcement Learning Approach in Multi-Phase Second-Price Auction Design
di: Ai, Rui, et al.
Pubblicazione: (2022)
di: Ai, Rui, et al.
Pubblicazione: (2022)
Incentive-Theoretic Bayesian Inference for Collaborative Science
di: Bates, Stephen, et al.
Pubblicazione: (2023)
di: Bates, Stephen, et al.
Pubblicazione: (2023)
Learning to Mitigate Externalities: the Coase Theorem with Hindsight Rationality
di: Scheid, Antoine, et al.
Pubblicazione: (2024)
di: Scheid, Antoine, et al.
Pubblicazione: (2024)
Last-Iterate Convergence of Payoff-Based Independent Learning in Zero-Sum Stochastic Games
di: Chen, Zaiwei, et al.
Pubblicazione: (2024)
di: Chen, Zaiwei, et al.
Pubblicazione: (2024)
Learning Dynamic Mechanisms in Unknown Environments: A Reinforcement Learning Approach
di: Qiu, Shuang, et al.
Pubblicazione: (2022)
di: Qiu, Shuang, et al.
Pubblicazione: (2022)
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria
di: Yu, Eason, et al.
Pubblicazione: (2025)
di: Yu, Eason, et al.
Pubblicazione: (2025)
Is Thompson Sampling Susceptible to Algorithmic Collusion?
di: Xiong, Yi, et al.
Pubblicazione: (2024)
di: Xiong, Yi, et al.
Pubblicazione: (2024)
Last-Iterate Convergence Properties of Regret-Matching Algorithms in Games
di: Cai, Yang, et al.
Pubblicazione: (2023)
di: Cai, Yang, et al.
Pubblicazione: (2023)
Last-Iterate Convergence of No-Regret Learning for Equilibria in Bargaining Games
di: Kamp, Serafina, et al.
Pubblicazione: (2025)
di: Kamp, Serafina, et al.
Pubblicazione: (2025)
Adaptive, Doubly Optimal No-Regret Learning in Strongly Monotone and Exp-Concave Games with Gradient Feedback
di: Jordan, Michael I., et al.
Pubblicazione: (2023)
di: Jordan, Michael I., et al.
Pubblicazione: (2023)
Efficient Last-Iterate Convergence in Regret Minimization via Adaptive Reward Transformation
di: Ren, Hang, et al.
Pubblicazione: (2025)
di: Ren, Hang, et al.
Pubblicazione: (2025)
Strategic Filtering for Content Moderation: Free Speech or Free of Distortion?
di: Ahmadi, Saba, et al.
Pubblicazione: (2025)
di: Ahmadi, Saba, et al.
Pubblicazione: (2025)
Efficient Uncoupled Learning Dynamics with $\tilde{O}\!\left(T^{-1/4}\right)$ Last-Iterate Convergence in Bilinear Saddle-Point Problems over Convex Sets under Bandit Feedback
di: Maiti, Arnab, et al.
Pubblicazione: (2026)
di: Maiti, Arnab, et al.
Pubblicazione: (2026)
Incentivized Learning in Principal-Agent Bandit Games
di: Scheid, Antoine, et al.
Pubblicazione: (2024)
di: Scheid, Antoine, et al.
Pubblicazione: (2024)
Independent Learning of Nash Equilibria in Partially Observable Markov Potential Games with Decoupled Dynamics
di: Jordan, Philip, et al.
Pubblicazione: (2026)
di: Jordan, Philip, et al.
Pubblicazione: (2026)
On Separation Between Best-Iterate, Random-Iterate, and Last-Iterate Convergence of Learning in Games
di: Cai, Yang, et al.
Pubblicazione: (2025)
di: Cai, Yang, et al.
Pubblicazione: (2025)
Towards Sustainable Investment Policies Informed by Opponent Shaping
di: Duque, Juan Agustin, et al.
Pubblicazione: (2026)
di: Duque, Juan Agustin, et al.
Pubblicazione: (2026)
Distributional Alignment Games for Answer-Level Fine-Tuning
di: Mohri, Mehryar, et al.
Pubblicazione: (2026)
di: Mohri, Mehryar, et al.
Pubblicazione: (2026)
Incentivizing High-Quality Content in Online Recommender Systems
di: Hu, Xinyan, et al.
Pubblicazione: (2023)
di: Hu, Xinyan, et al.
Pubblicazione: (2023)
Improved Bayes Risk Can Yield Reduced Social Welfare Under Competition
di: Jagadeesan, Meena, et al.
Pubblicazione: (2023)
di: Jagadeesan, Meena, et al.
Pubblicazione: (2023)
Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?
di: Gölz, Paul, et al.
Pubblicazione: (2025)
di: Gölz, Paul, et al.
Pubblicazione: (2025)
Independent Learning in Constrained Markov Potential Games
di: Jordan, Philip, et al.
Pubblicazione: (2024)
di: Jordan, Philip, et al.
Pubblicazione: (2024)
How to Evaluate Behavioral Models
di: d'Eon, Greg, et al.
Pubblicazione: (2023)
di: d'Eon, Greg, et al.
Pubblicazione: (2023)
State-Constrained Zero-Sum Differential Games with One-Sided Information
di: Ghimire, Mukesh, et al.
Pubblicazione: (2024)
di: Ghimire, Mukesh, et al.
Pubblicazione: (2024)
Analysing the Sample Complexity of Opponent Shaping
di: Fung, Kitty, et al.
Pubblicazione: (2024)
di: Fung, Kitty, et al.
Pubblicazione: (2024)
Clickbait vs. Quality: How Engagement-Based Optimization Shapes the Content Landscape in Online Platforms
di: Immorlica, Nicole, et al.
Pubblicazione: (2024)
di: Immorlica, Nicole, et al.
Pubblicazione: (2024)
Are Bounded Contracts Learnable and Approximately Optimal?
di: Chen, Yurong, et al.
Pubblicazione: (2024)
di: Chen, Yurong, et al.
Pubblicazione: (2024)
How Market Volatility Shapes Algorithmic Collusion: A Comparative Analysis of Learning-Based Pricing Algorithms
di: Sravon, Aheer, et al.
Pubblicazione: (2025)
di: Sravon, Aheer, et al.
Pubblicazione: (2025)
Refined Sample Complexity for Markov Games with Independent Linear Function Approximation
di: Dai, Yan, et al.
Pubblicazione: (2024)
di: Dai, Yan, et al.
Pubblicazione: (2024)
Sample Efficient Omniprediction and Downstream Swap Regret for Non-Linear Losses
di: Lu, Jiuyao, et al.
Pubblicazione: (2025)
di: Lu, Jiuyao, et al.
Pubblicazione: (2025)
Geometry Meets Incentives: Sample-Efficient Incentivized Exploration with Linear Contexts
di: Schiffer, Benjamin, et al.
Pubblicazione: (2025)
di: Schiffer, Benjamin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Calibeating Made Simple
di: Chen, Yurong, et al.
Pubblicazione: (2026) -
Response Time Enhances Alignment with Heterogeneous Preferences
di: Echenique, Federico, et al.
Pubblicazione: (2026) -
One-Shot Strategic Classification Under Unknown Costs
di: Rosenfeld, Elan, et al.
Pubblicazione: (2023) -
From Average-Iterate to Last-Iterate Convergence in Games: A Reduction and Its Applications
di: Cai, Yang, et al.
Pubblicazione: (2025) -
Last-Iterate Convergence of Adaptive Riemannian Gradient Descent for Equilibrium Computation
di: Cai, Yang, et al.
Pubblicazione: (2023)