Online SuBmodular + SuPermodular (BP) Maximization with Bandit Feedback
Fuente:
arXiv
Guardado en:
| Autores principales: | Narang, Adhyyan, Sadeghi, Omid, Ratliff, Lillian J, Fazel, Maryam, Bilmes, Jeff |
|---|---|
| Formato: | Preprint |
| Publicado: |
2022
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Dynamics of Learning under User Choice: Overspecialization and Peer-Model Probing
por: Narang, Adhyyan, et al.
Publicado: (2026)
por: Narang, Adhyyan, et al.
Publicado: (2026)
Sample Complexity Reduction via Policy Difference Estimation in Tabular Reinforcement Learning
por: Narang, Adhyyan, et al.
Publicado: (2024)
por: Narang, Adhyyan, et al.
Publicado: (2024)
Principal-Agent Bandit Games with Self-Interested and Exploratory Learning Agents
por: Liu, Junyan, et al.
Publicado: (2024)
por: Liu, Junyan, et al.
Publicado: (2024)
Efficient Near-Optimal Algorithm for Online Shortest Paths in Directed Acyclic Graphs with Bandit Feedback Against Adaptive Adversaries
por: Maiti, Arnab, et al.
Publicado: (2025)
por: Maiti, Arnab, et al.
Publicado: (2025)
Emergent specialization from participation dynamics and multi-learner retraining
por: Dean, Sarah, et al.
Publicado: (2022)
por: Dean, Sarah, et al.
Publicado: (2022)
On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback
por: Williams, Marcus, et al.
Publicado: (2024)
por: Williams, Marcus, et al.
Publicado: (2024)
On the Limitations and Possibilities of Nash Regret Minimization in Zero-Sum Matrix Games under Noisy Feedback
por: Maiti, Arnab, et al.
Publicado: (2023)
por: Maiti, Arnab, et al.
Publicado: (2023)
Initializing Services in Interactive ML Systems for Diverse Users
por: Bose, Avinandan, et al.
Publicado: (2023)
por: Bose, Avinandan, et al.
Publicado: (2023)
On The Complexity of Best-Arm Identification in Non-Stationary Linear Bandits
por: Maynard-Zhang, Leo, et al.
Publicado: (2026)
por: Maynard-Zhang, Leo, et al.
Publicado: (2026)
Extragradient Preference Optimization (EGPO): Beyond Last-Iterate Convergence for Nash Learning from Human Feedback
por: Zhou, Runlong, et al.
Publicado: (2025)
por: Zhou, Runlong, et al.
Publicado: (2025)
Adaptive Calibration in Non-Stationary Environments
por: Liu, Junyan, et al.
Publicado: (2026)
por: Liu, Junyan, et al.
Publicado: (2026)
Linear Submodular Maximization with Bandit Feedback
por: Chen, Wenjing, et al.
Publicado: (2024)
por: Chen, Wenjing, et al.
Publicado: (2024)
Online Learning for Uninformed Markov Games: Empirical Nash-Value Regret and Non-Stationarity Adaptation
por: Liu, Junyan, et al.
Publicado: (2026)
por: Liu, Junyan, et al.
Publicado: (2026)
Nearly Minimax Optimal Submodular Maximization with Bandit Feedback
por: Tajdini, Artin, et al.
Publicado: (2023)
por: Tajdini, Artin, et al.
Publicado: (2023)
Efficient Uncoupled Learning Dynamics with $\tilde{O}\!\left(T^{-1/4}\right)$ Last-Iterate Convergence in Bilinear Saddle-Point Problems over Convex Sets under Bandit Feedback
por: Maiti, Arnab, et al.
Publicado: (2026)
por: Maiti, Arnab, et al.
Publicado: (2026)
A/B Testing and Best-arm Identification for Linear Bandits with Robustness to Non-stationarity
por: Xiong, Zhihan, et al.
Publicado: (2023)
por: Xiong, Zhihan, et al.
Publicado: (2023)
Beyond Bandit Feedback in Online Multiclass Classification
por: van der Hoeven, Dirk, et al.
Publicado: (2021)
por: van der Hoeven, Dirk, et al.
Publicado: (2021)
Bandit and Delayed Feedback in Online Structured Prediction
por: Shibukawa, Yuki, et al.
Publicado: (2025)
por: Shibukawa, Yuki, et al.
Publicado: (2025)
Multiclass Online Learnability under Bandit Feedback
por: Raman, Ananth, et al.
Publicado: (2023)
por: Raman, Ananth, et al.
Publicado: (2023)
Explore-then-Commit for Nonstationary Linear Bandits with Latent Dynamics
por: Choi, Sunmook, et al.
Publicado: (2025)
por: Choi, Sunmook, et al.
Publicado: (2025)
Deep Submodular Peripteral Networks
por: Bhatt, Gantavya, et al.
Publicado: (2024)
por: Bhatt, Gantavya, et al.
Publicado: (2024)
Unregularized Linear Convergence in Zero-Sum Game from Preference Feedback
por: Chen, Shulun, et al.
Publicado: (2025)
por: Chen, Shulun, et al.
Publicado: (2025)
Stochastic Online Conformal Prediction with Semi-Bandit Feedback
por: Ge, Haosen, et al.
Publicado: (2024)
por: Ge, Haosen, et al.
Publicado: (2024)
Efficient Online Set-valued Classification with Bandit Feedback
por: Wang, Zhou, et al.
Publicado: (2024)
por: Wang, Zhou, et al.
Publicado: (2024)
Bandit-Feedback Online Multiclass Classification: Variants and Tradeoffs
por: Filmus, Yuval, et al.
Publicado: (2024)
por: Filmus, Yuval, et al.
Publicado: (2024)
Online Nonsubmodular Optimization with Delayed Feedback in the Bandit Setting
por: Yang, Sifan, et al.
Publicado: (2025)
por: Yang, Sifan, et al.
Publicado: (2025)
dUltra: Ultra-Fast Diffusion Language Models via Reinforcement Learning
por: Chen, Shirui, et al.
Publicado: (2025)
por: Chen, Shirui, et al.
Publicado: (2025)
Convergence of Learning Dynamics in Stackelberg Games
por: Fiez, Tanner, et al.
Publicado: (2019)
por: Fiez, Tanner, et al.
Publicado: (2019)
Learning to Schedule Online Tasks with Bandit Feedback
por: Xu, Yongxin, et al.
Publicado: (2024)
por: Xu, Yongxin, et al.
Publicado: (2024)
Stochastic Online Instrumental Variable Regression: Regrets for Endogeneity and Bandit Feedback
por: Della Vecchia, Riccardo, et al.
Publicado: (2023)
por: Della Vecchia, Riccardo, et al.
Publicado: (2023)
Online Conformal Abstention for Factuality Control Under Adversarial Bandit Feedback
por: Lee, Minjae, et al.
Publicado: (2025)
por: Lee, Minjae, et al.
Publicado: (2025)
Dual Approximation Policy Optimization
por: Xiong, Zhihan, et al.
Publicado: (2024)
por: Xiong, Zhihan, et al.
Publicado: (2024)
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration
por: Bose, Avinandan, et al.
Publicado: (2024)
por: Bose, Avinandan, et al.
Publicado: (2024)
A Learning Algorithm That Attains the Human Optimum in a Repeated Human-Machine Interaction Game
por: Isa, Jason T., et al.
Publicado: (2025)
por: Isa, Jason T., et al.
Publicado: (2025)
Adaptive Client Sampling in Federated Learning via Online Learning with Bandit Feedback
por: Zhao, Boxin, et al.
Publicado: (2021)
por: Zhao, Boxin, et al.
Publicado: (2021)
Online Budget Allocation with Censored Semi-Bandit Feedback
por: Bachoc, François, et al.
Publicado: (2025)
por: Bachoc, François, et al.
Publicado: (2025)
Many-Objective Multi-Solution Transport
por: Li, Ziyue, et al.
Publicado: (2024)
por: Li, Ziyue, et al.
Publicado: (2024)
Strategically Robust Multi-Agent Reinforcement Learning with Linear Function Approximation
por: Gonzales, Jake, et al.
Publicado: (2026)
por: Gonzales, Jake, et al.
Publicado: (2026)
DOPL: Direct Online Preference Learning for Restless Bandits with Preference Feedback
por: Xiong, Guojun, et al.
Publicado: (2024)
por: Xiong, Guojun, et al.
Publicado: (2024)
Near-optimal Regret Using Policy Optimization in Online MDPs with Aggregate Bandit Feedback
por: Lancewicki, Tal, et al.
Publicado: (2025)
por: Lancewicki, Tal, et al.
Publicado: (2025)
Ejemplares similares
-
Dynamics of Learning under User Choice: Overspecialization and Peer-Model Probing
por: Narang, Adhyyan, et al.
Publicado: (2026) -
Sample Complexity Reduction via Policy Difference Estimation in Tabular Reinforcement Learning
por: Narang, Adhyyan, et al.
Publicado: (2024) -
Principal-Agent Bandit Games with Self-Interested and Exploratory Learning Agents
por: Liu, Junyan, et al.
Publicado: (2024) -
Efficient Near-Optimal Algorithm for Online Shortest Paths in Directed Acyclic Graphs with Bandit Feedback Against Adaptive Adversaries
por: Maiti, Arnab, et al.
Publicado: (2025) -
Emergent specialization from participation dynamics and multi-learner retraining
por: Dean, Sarah, et al.
Publicado: (2022)