Online SuBmodular + SuPermodular (BP) Maximization with Bandit Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | Narang, Adhyyan, Sadeghi, Omid, Ratliff, Lillian J, Fazel, Maryam, Bilmes, Jeff |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dynamics of Learning under User Choice: Overspecialization and Peer-Model Probing
by: Narang, Adhyyan, et al.
Published: (2026)
by: Narang, Adhyyan, et al.
Published: (2026)
Sample Complexity Reduction via Policy Difference Estimation in Tabular Reinforcement Learning
by: Narang, Adhyyan, et al.
Published: (2024)
by: Narang, Adhyyan, et al.
Published: (2024)
Principal-Agent Bandit Games with Self-Interested and Exploratory Learning Agents
by: Liu, Junyan, et al.
Published: (2024)
by: Liu, Junyan, et al.
Published: (2024)
Efficient Near-Optimal Algorithm for Online Shortest Paths in Directed Acyclic Graphs with Bandit Feedback Against Adaptive Adversaries
by: Maiti, Arnab, et al.
Published: (2025)
by: Maiti, Arnab, et al.
Published: (2025)
Emergent specialization from participation dynamics and multi-learner retraining
by: Dean, Sarah, et al.
Published: (2022)
by: Dean, Sarah, et al.
Published: (2022)
On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback
by: Williams, Marcus, et al.
Published: (2024)
by: Williams, Marcus, et al.
Published: (2024)
On the Limitations and Possibilities of Nash Regret Minimization in Zero-Sum Matrix Games under Noisy Feedback
by: Maiti, Arnab, et al.
Published: (2023)
by: Maiti, Arnab, et al.
Published: (2023)
Initializing Services in Interactive ML Systems for Diverse Users
by: Bose, Avinandan, et al.
Published: (2023)
by: Bose, Avinandan, et al.
Published: (2023)
On The Complexity of Best-Arm Identification in Non-Stationary Linear Bandits
by: Maynard-Zhang, Leo, et al.
Published: (2026)
by: Maynard-Zhang, Leo, et al.
Published: (2026)
Extragradient Preference Optimization (EGPO): Beyond Last-Iterate Convergence for Nash Learning from Human Feedback
by: Zhou, Runlong, et al.
Published: (2025)
by: Zhou, Runlong, et al.
Published: (2025)
Adaptive Calibration in Non-Stationary Environments
by: Liu, Junyan, et al.
Published: (2026)
by: Liu, Junyan, et al.
Published: (2026)
Linear Submodular Maximization with Bandit Feedback
by: Chen, Wenjing, et al.
Published: (2024)
by: Chen, Wenjing, et al.
Published: (2024)
Online Learning for Uninformed Markov Games: Empirical Nash-Value Regret and Non-Stationarity Adaptation
by: Liu, Junyan, et al.
Published: (2026)
by: Liu, Junyan, et al.
Published: (2026)
Nearly Minimax Optimal Submodular Maximization with Bandit Feedback
by: Tajdini, Artin, et al.
Published: (2023)
by: Tajdini, Artin, et al.
Published: (2023)
Efficient Uncoupled Learning Dynamics with $\tilde{O}\!\left(T^{-1/4}\right)$ Last-Iterate Convergence in Bilinear Saddle-Point Problems over Convex Sets under Bandit Feedback
by: Maiti, Arnab, et al.
Published: (2026)
by: Maiti, Arnab, et al.
Published: (2026)
A/B Testing and Best-arm Identification for Linear Bandits with Robustness to Non-stationarity
by: Xiong, Zhihan, et al.
Published: (2023)
by: Xiong, Zhihan, et al.
Published: (2023)
Beyond Bandit Feedback in Online Multiclass Classification
by: van der Hoeven, Dirk, et al.
Published: (2021)
by: van der Hoeven, Dirk, et al.
Published: (2021)
Bandit and Delayed Feedback in Online Structured Prediction
by: Shibukawa, Yuki, et al.
Published: (2025)
by: Shibukawa, Yuki, et al.
Published: (2025)
Multiclass Online Learnability under Bandit Feedback
by: Raman, Ananth, et al.
Published: (2023)
by: Raman, Ananth, et al.
Published: (2023)
Explore-then-Commit for Nonstationary Linear Bandits with Latent Dynamics
by: Choi, Sunmook, et al.
Published: (2025)
by: Choi, Sunmook, et al.
Published: (2025)
Deep Submodular Peripteral Networks
by: Bhatt, Gantavya, et al.
Published: (2024)
by: Bhatt, Gantavya, et al.
Published: (2024)
Unregularized Linear Convergence in Zero-Sum Game from Preference Feedback
by: Chen, Shulun, et al.
Published: (2025)
by: Chen, Shulun, et al.
Published: (2025)
Stochastic Online Conformal Prediction with Semi-Bandit Feedback
by: Ge, Haosen, et al.
Published: (2024)
by: Ge, Haosen, et al.
Published: (2024)
Efficient Online Set-valued Classification with Bandit Feedback
by: Wang, Zhou, et al.
Published: (2024)
by: Wang, Zhou, et al.
Published: (2024)
Bandit-Feedback Online Multiclass Classification: Variants and Tradeoffs
by: Filmus, Yuval, et al.
Published: (2024)
by: Filmus, Yuval, et al.
Published: (2024)
Online Nonsubmodular Optimization with Delayed Feedback in the Bandit Setting
by: Yang, Sifan, et al.
Published: (2025)
by: Yang, Sifan, et al.
Published: (2025)
dUltra: Ultra-Fast Diffusion Language Models via Reinforcement Learning
by: Chen, Shirui, et al.
Published: (2025)
by: Chen, Shirui, et al.
Published: (2025)
Convergence of Learning Dynamics in Stackelberg Games
by: Fiez, Tanner, et al.
Published: (2019)
by: Fiez, Tanner, et al.
Published: (2019)
Learning to Schedule Online Tasks with Bandit Feedback
by: Xu, Yongxin, et al.
Published: (2024)
by: Xu, Yongxin, et al.
Published: (2024)
Stochastic Online Instrumental Variable Regression: Regrets for Endogeneity and Bandit Feedback
by: Della Vecchia, Riccardo, et al.
Published: (2023)
by: Della Vecchia, Riccardo, et al.
Published: (2023)
Online Conformal Abstention for Factuality Control Under Adversarial Bandit Feedback
by: Lee, Minjae, et al.
Published: (2025)
by: Lee, Minjae, et al.
Published: (2025)
Dual Approximation Policy Optimization
by: Xiong, Zhihan, et al.
Published: (2024)
by: Xiong, Zhihan, et al.
Published: (2024)
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration
by: Bose, Avinandan, et al.
Published: (2024)
by: Bose, Avinandan, et al.
Published: (2024)
A Learning Algorithm That Attains the Human Optimum in a Repeated Human-Machine Interaction Game
by: Isa, Jason T., et al.
Published: (2025)
by: Isa, Jason T., et al.
Published: (2025)
Adaptive Client Sampling in Federated Learning via Online Learning with Bandit Feedback
by: Zhao, Boxin, et al.
Published: (2021)
by: Zhao, Boxin, et al.
Published: (2021)
Online Budget Allocation with Censored Semi-Bandit Feedback
by: Bachoc, François, et al.
Published: (2025)
by: Bachoc, François, et al.
Published: (2025)
Many-Objective Multi-Solution Transport
by: Li, Ziyue, et al.
Published: (2024)
by: Li, Ziyue, et al.
Published: (2024)
Strategically Robust Multi-Agent Reinforcement Learning with Linear Function Approximation
by: Gonzales, Jake, et al.
Published: (2026)
by: Gonzales, Jake, et al.
Published: (2026)
DOPL: Direct Online Preference Learning for Restless Bandits with Preference Feedback
by: Xiong, Guojun, et al.
Published: (2024)
by: Xiong, Guojun, et al.
Published: (2024)
Near-optimal Regret Using Policy Optimization in Online MDPs with Aggregate Bandit Feedback
by: Lancewicki, Tal, et al.
Published: (2025)
by: Lancewicki, Tal, et al.
Published: (2025)
Similar Items
-
Dynamics of Learning under User Choice: Overspecialization and Peer-Model Probing
by: Narang, Adhyyan, et al.
Published: (2026) -
Sample Complexity Reduction via Policy Difference Estimation in Tabular Reinforcement Learning
by: Narang, Adhyyan, et al.
Published: (2024) -
Principal-Agent Bandit Games with Self-Interested and Exploratory Learning Agents
by: Liu, Junyan, et al.
Published: (2024) -
Efficient Near-Optimal Algorithm for Online Shortest Paths in Directed Acyclic Graphs with Bandit Feedback Against Adaptive Adversaries
by: Maiti, Arnab, et al.
Published: (2025) -
Emergent specialization from participation dynamics and multi-learner retraining
by: Dean, Sarah, et al.
Published: (2022)