Bandits with Preference Feedback: A Stackelberg Game Perspective
Fuente:
arXiv
Guardado en:
| Autores principales: | Pásztor, Barna, Kassraie, Parnian, Krause, Andreas |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Stackelberg Learning from Human Feedback: Preference Optimization as a Sequential Game
por: Pásztor, Barna, et al.
Publicado: (2025)
por: Pásztor, Barna, et al.
Publicado: (2025)
Self-optimization in distributed manufacturing systems using Modular State-based Stackelberg Games
por: Yuwono, Steve, et al.
Publicado: (2024)
por: Yuwono, Steve, et al.
Publicado: (2024)
The Intelligent Disobedience Game: Formulating Disobedience in Stackelberg Games and Markov Decision Processes
por: Hornig, Benedikt, et al.
Publicado: (2026)
por: Hornig, Benedikt, et al.
Publicado: (2026)
Model-Based RL for Mean-Field Games is not Statistically Harder than Single-Agent RL
por: Huang, Jiawei, et al.
Publicado: (2024)
por: Huang, Jiawei, et al.
Publicado: (2024)
Preferences Evolve And So Should Your Bandits: Bandits with Evolving States for Online Platforms
por: Khosravi, Khashayar, et al.
Publicado: (2023)
por: Khosravi, Khashayar, et al.
Publicado: (2023)
Nearly-Optimal Bandit Learning in Stackelberg Games with Side Information
por: Balcan, Maria-Florina, et al.
Publicado: (2025)
por: Balcan, Maria-Florina, et al.
Publicado: (2025)
Distributed Stackelberg Strategies in State-based Potential Games for Autonomous Decentralized Learning Manufacturing Systems
por: Yuwono, Steve, et al.
Publicado: (2024)
por: Yuwono, Steve, et al.
Publicado: (2024)
Randomised Optimism via Competitive Co-Evolution for Matrix Games with Bandit Feedback
por: Lin, Shishen
Publicado: (2025)
por: Lin, Shishen
Publicado: (2025)
Meta-Computing Enhanced Federated Learning in IIoT: Satisfaction-Aware Incentive Scheme via DRL-Based Stackelberg Game
por: Li, Xiaohuan, et al.
Publicado: (2025)
por: Li, Xiaohuan, et al.
Publicado: (2025)
A Resilience Framework for Bi-Criteria Combinatorial Optimization with Bandit Feedback
por: Aggarwal, Vaneet, et al.
Publicado: (2025)
por: Aggarwal, Vaneet, et al.
Publicado: (2025)
Two-Player Zero-Sum Games with Bandit Feedback
por: Yılmaz, Elif, et al.
Publicado: (2025)
por: Yılmaz, Elif, et al.
Publicado: (2025)
Learning in Structured Stackelberg Games
por: Balcan, Maria-Florina, et al.
Publicado: (2025)
por: Balcan, Maria-Florina, et al.
Publicado: (2025)
RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models
por: Yang, Daniel, et al.
Publicado: (2026)
por: Yang, Daniel, et al.
Publicado: (2026)
Gradient-based Learning in State-based Potential Games for Self-Learning Production Systems
por: Yuwono, Steve, et al.
Publicado: (2024)
por: Yuwono, Steve, et al.
Publicado: (2024)
Robust Solutions for Multi-Defender Stackelberg Security Games
por: Mutzari, Dolev, et al.
Publicado: (2022)
por: Mutzari, Dolev, et al.
Publicado: (2022)
Regret Minimization in Stackelberg Games with Side Information
por: Harris, Keegan, et al.
Publicado: (2024)
por: Harris, Keegan, et al.
Publicado: (2024)
Learning in Repeated Multi-Objective Stackelberg Games with Payoff Manipulation
por: Srisawad, Phurinut, et al.
Publicado: (2025)
por: Srisawad, Phurinut, et al.
Publicado: (2025)
Paths to Equilibrium in Games
por: Yongacoglu, Bora, et al.
Publicado: (2024)
por: Yongacoglu, Bora, et al.
Publicado: (2024)
The Hidden Game Problem
por: Buzaglo, Gon, et al.
Publicado: (2025)
por: Buzaglo, Gon, et al.
Publicado: (2025)
Impact of Decentralized Learning on Player Utilities in Stackelberg Games
por: Donahue, Kate, et al.
Publicado: (2024)
por: Donahue, Kate, et al.
Publicado: (2024)
Zeroth-Order Stackelberg Control in Combinatorial Congestion Games
por: Masiha, Saeed, et al.
Publicado: (2026)
por: Masiha, Saeed, et al.
Publicado: (2026)
Learning in Bayesian Stackelberg Games With Unknown Follower's Types
por: Bollini, Matteo, et al.
Publicado: (2026)
por: Bollini, Matteo, et al.
Publicado: (2026)
LLM-Powered Preference Elicitation in Combinatorial Assignment
por: Soumalias, Ermis, et al.
Publicado: (2025)
por: Soumalias, Ermis, et al.
Publicado: (2025)
Axioms for AI Alignment from Human Feedback
por: Ge, Luise, et al.
Publicado: (2024)
por: Ge, Luise, et al.
Publicado: (2024)
Efficient Ensemble Selection from Binary and Pairwise Feedback
por: Neoh, Tzeh Yuan, et al.
Publicado: (2026)
por: Neoh, Tzeh Yuan, et al.
Publicado: (2026)
Monopoly Deal: A Benchmark Environment for Bounded One-Sided Response Games
por: Wolf, Will
Publicado: (2025)
por: Wolf, Will
Publicado: (2025)
Optimistic Thompson Sampling for No-Regret Learning in Unknown Games
por: Li, Yingru, et al.
Publicado: (2024)
por: Li, Yingru, et al.
Publicado: (2024)
Efficient Last-iterate Convergence Algorithms in Solving Games
por: Meng, Linjian, et al.
Publicado: (2023)
por: Meng, Linjian, et al.
Publicado: (2023)
A Policy-Gradient Approach to Solving Imperfect-Information Games with Best-Iterate Convergence
por: Liu, Mingyang, et al.
Publicado: (2024)
por: Liu, Mingyang, et al.
Publicado: (2024)
Do LLM Agents Have Regret? A Case Study in Online Learning and Games
por: Park, Chanwoo, et al.
Publicado: (2024)
por: Park, Chanwoo, et al.
Publicado: (2024)
Regret Minimization and Convergence to Equilibria in General-sum Markov Games
por: Erez, Liad, et al.
Publicado: (2022)
por: Erez, Liad, et al.
Publicado: (2022)
Incentivizing Truthful Language Models via Peer Elicitation Games
por: Chen, Baiting, et al.
Publicado: (2025)
por: Chen, Baiting, et al.
Publicado: (2025)
Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers
por: Sun, Haoran, et al.
Publicado: (2025)
por: Sun, Haoran, et al.
Publicado: (2025)
LiteEFG: An Efficient Python Library for Solving Extensive-form Games
por: Liu, Mingyang, et al.
Publicado: (2024)
por: Liu, Mingyang, et al.
Publicado: (2024)
When Should a Leader Act Suboptimally? The Role of Inferability in Repeated Stackelberg Games
por: Karabag, Mustafa O., et al.
Publicado: (2023)
por: Karabag, Mustafa O., et al.
Publicado: (2023)
ElementaryNet: A Non-Strategic Neural Network for Predicting Human Behavior in Normal-Form Games
por: d'Eon, Greg, et al.
Publicado: (2025)
por: d'Eon, Greg, et al.
Publicado: (2025)
Do LLMs Strategically Reveal, Conceal, and Infer Information? A Theoretical and Empirical Analysis in The Chameleon Game
por: Karabag, Mustafa O., et al.
Publicado: (2025)
por: Karabag, Mustafa O., et al.
Publicado: (2025)
Efficiently Training Neural Networks for Imperfect Information Games by Sampling Information Sets
por: Bertram, Timo, et al.
Publicado: (2024)
por: Bertram, Timo, et al.
Publicado: (2024)
Online Test Synthesis From Requirements: Enhancing Reinforcement Learning with Game Theory
por: Sankur, Ocan, et al.
Publicado: (2024)
por: Sankur, Ocan, et al.
Publicado: (2024)
Improving Sample Efficiency of Model-Free Algorithms for Zero-Sum Markov Games
por: Feng, Songtao, et al.
Publicado: (2023)
por: Feng, Songtao, et al.
Publicado: (2023)
Ejemplares similares
-
Stackelberg Learning from Human Feedback: Preference Optimization as a Sequential Game
por: Pásztor, Barna, et al.
Publicado: (2025) -
Self-optimization in distributed manufacturing systems using Modular State-based Stackelberg Games
por: Yuwono, Steve, et al.
Publicado: (2024) -
The Intelligent Disobedience Game: Formulating Disobedience in Stackelberg Games and Markov Decision Processes
por: Hornig, Benedikt, et al.
Publicado: (2026) -
Model-Based RL for Mean-Field Games is not Statistically Harder than Single-Agent RL
por: Huang, Jiawei, et al.
Publicado: (2024) -
Preferences Evolve And So Should Your Bandits: Bandits with Evolving States for Online Platforms
por: Khosravi, Khashayar, et al.
Publicado: (2023)