Efficient Contextual Bandits with Uninformed Feedback Graphs
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Mengxiao, Zhang, Yuheng, Luo, Haipeng, Mineiro, Paul |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Interaction-Grounded Learning for Contextual Markov Decision Processes with Personalized Feedback
por: Zhang, Mengxiao, et al.
Publicado: (2026)
por: Zhang, Mengxiao, et al.
Publicado: (2026)
Provably Efficient Interactive-Grounded Learning with Personalized Reward
por: Zhang, Mengxiao, et al.
Publicado: (2024)
por: Zhang, Mengxiao, et al.
Publicado: (2024)
Contextual Multinomial Logit Bandits with General Value Functions
por: Zhang, Mengxiao, et al.
Publicado: (2024)
por: Zhang, Mengxiao, et al.
Publicado: (2024)
Contextual Linear Bandits with Delay as Payoff
por: Zhang, Mengxiao, et al.
Publicado: (2025)
por: Zhang, Mengxiao, et al.
Publicado: (2025)
Near-Optimal Last-Iterate Convergence for Zero-Sum Games with Bandit Feedback and Opponent Actions
por: Hait, Soumita, et al.
Publicado: (2026)
por: Hait, Soumita, et al.
Publicado: (2026)
Efficient Algorithms for Logistic Contextual Slate Bandits with Bandit Feedback
por: Goyal, Tanmay, et al.
Publicado: (2025)
por: Goyal, Tanmay, et al.
Publicado: (2025)
Online Learning for Uninformed Markov Games: Empirical Nash-Value Regret and Non-Stationarity Adaptation
por: Liu, Junyan, et al.
Publicado: (2026)
por: Liu, Junyan, et al.
Publicado: (2026)
Near-Optimal Regret in Linear MDPs with Aggregate Bandit Feedback
por: Cassel, Asaf, et al.
Publicado: (2024)
por: Cassel, Asaf, et al.
Publicado: (2024)
Alternating Regret for Online Convex Optimization
por: Hait, Soumita, et al.
Publicado: (2025)
por: Hait, Soumita, et al.
Publicado: (2025)
Comparator-Adaptive $Φ$-Regret: Improved Bounds, Simpler Algorithms, and Applications to Games
por: Hait, Soumita, et al.
Publicado: (2025)
por: Hait, Soumita, et al.
Publicado: (2025)
Adapting to Stochastic and Adversarial Losses in Episodic MDPs with Aggregate Bandit Feedback
por: Ito, Shinji, et al.
Publicado: (2025)
por: Ito, Shinji, et al.
Publicado: (2025)
Online Joint Fine-tuning of Multi-Agent Flows
por: Mineiro, Paul
Publicado: (2024)
por: Mineiro, Paul
Publicado: (2024)
Neural Contextual Bandits Under Delayed Feedback Constraints
por: Moghimi, Mohammadali, et al.
Publicado: (2025)
por: Moghimi, Mohammadali, et al.
Publicado: (2025)
Locally Private Nonparametric Contextual Multi-armed Bandits
por: Ma, Yuheng, et al.
Publicado: (2025)
por: Ma, Yuheng, et al.
Publicado: (2025)
Graph Feedback Bandits on Similar Arms: With and Without Graph Structures
por: Qi, Han, et al.
Publicado: (2025)
por: Qi, Han, et al.
Publicado: (2025)
Near-Optimal Regret for Distributed Adversarial Bandits: A Black-Box Approach
por: Qiu, Hao, et al.
Publicado: (2026)
por: Qiu, Hao, et al.
Publicado: (2026)
Exploiting Curvature in Online Convex Optimization with Delayed Feedback
por: Qiu, Hao, et al.
Publicado: (2025)
por: Qiu, Hao, et al.
Publicado: (2025)
No-Regret Learning for Fair Multi-Agent Social Welfare Optimization
por: Zhang, Mengxiao, et al.
Publicado: (2024)
por: Zhang, Mengxiao, et al.
Publicado: (2024)
Sparse Nonparametric Contextual Bandits
por: Flynn, Hamish, et al.
Publicado: (2025)
por: Flynn, Hamish, et al.
Publicado: (2025)
IBCB: Efficient Inverse Batched Contextual Bandit for Behavioral Evolution History
por: Xu, Yi, et al.
Publicado: (2024)
por: Xu, Yi, et al.
Publicado: (2024)
Graph Feedback Bandits with Similar Arms
por: Qi, Han, et al.
Publicado: (2024)
por: Qi, Han, et al.
Publicado: (2024)
Nearly Optimal Algorithms for Contextual Dueling Bandits from Adversarial Feedback
por: Di, Qiwei, et al.
Publicado: (2024)
por: Di, Qiwei, et al.
Publicado: (2024)
Active Human Feedback Collection via Neural Contextual Dueling Bandits
por: Verma, Arun, et al.
Publicado: (2025)
por: Verma, Arun, et al.
Publicado: (2025)
Nearly Tight Bounds for Cross-Learning Contextual Bandits with Graphical Feedback
por: Huang, Ruiyuan, et al.
Publicado: (2025)
por: Huang, Ruiyuan, et al.
Publicado: (2025)
Decentralized Online Convex Optimization with Unknown Feedback Delays
por: Qiu, Hao, et al.
Publicado: (2026)
por: Qiu, Hao, et al.
Publicado: (2026)
Group-Sensitive Offline Contextual Bandits
por: Guo, Yihong, et al.
Publicado: (2025)
por: Guo, Yihong, et al.
Publicado: (2025)
Regret Bounds for Adversarial Contextual Bandits with General Function Approximation and Delayed Feedback
por: Levy, Orin, et al.
Publicado: (2025)
por: Levy, Orin, et al.
Publicado: (2025)
Efficient Generalized Low-Rank Tensor Contextual Bandits
por: Yi, Qianxin, et al.
Publicado: (2023)
por: Yi, Qianxin, et al.
Publicado: (2023)
Contextual Bandits for Unbounded Context Distributions
por: Zhao, Puning, et al.
Publicado: (2024)
por: Zhao, Puning, et al.
Publicado: (2024)
One Good Source is All You Need: Near-Optimal Regret for Bandits under Heterogeneous Noise
por: Bhat, Amith, et al.
Publicado: (2026)
por: Bhat, Amith, et al.
Publicado: (2026)
Generalized Low-Rank Matrix Contextual Bandits with Graph Information
por: Wang, Yao, et al.
Publicado: (2025)
por: Wang, Yao, et al.
Publicado: (2025)
Parameter-free Dynamic Regret: Time-varying Movement Costs, Delayed Feedback, and Memory
por: Qiu, Hao, et al.
Publicado: (2026)
por: Qiu, Hao, et al.
Publicado: (2026)
Taming the Monster Every Context: Complexity Measure and Unified Framework for Offline-Oracle Efficient Contextual Bandits
por: Qin, Hao, et al.
Publicado: (2026)
por: Qin, Hao, et al.
Publicado: (2026)
Single Index Bandits: Generalized Linear Contextual Bandits with Unknown Reward Functions
por: Kang, Yue, et al.
Publicado: (2025)
por: Kang, Yue, et al.
Publicado: (2025)
Flow-DPO: Improving LLM Mathematical Reasoning through Online Multi-Agent Learning
por: Deng, Yihe, et al.
Publicado: (2024)
por: Deng, Yihe, et al.
Publicado: (2024)
Recycling History: Efficient Recommendations from Contextual Dueling Bandits
por: Sankagiri, Suryanarayana, et al.
Publicado: (2025)
por: Sankagiri, Suryanarayana, et al.
Publicado: (2025)
Sharp Analysis for KL-Regularized Contextual Bandits and RLHF
por: Zhao, Heyang, et al.
Publicado: (2024)
por: Zhao, Heyang, et al.
Publicado: (2024)
Catoni Contextual Bandits are Robust to Heavy-tailed Rewards
por: Ye, Chenlu, et al.
Publicado: (2025)
por: Ye, Chenlu, et al.
Publicado: (2025)
Bi-Level Contextual Bandits for Individualized Resource Allocation under Delayed Feedback
por: Almasi, Mohammadsina, et al.
Publicado: (2025)
por: Almasi, Mohammadsina, et al.
Publicado: (2025)
Efficient Online Set-valued Classification with Bandit Feedback
por: Wang, Zhou, et al.
Publicado: (2024)
por: Wang, Zhou, et al.
Publicado: (2024)
Ejemplares similares
-
Interaction-Grounded Learning for Contextual Markov Decision Processes with Personalized Feedback
por: Zhang, Mengxiao, et al.
Publicado: (2026) -
Provably Efficient Interactive-Grounded Learning with Personalized Reward
por: Zhang, Mengxiao, et al.
Publicado: (2024) -
Contextual Multinomial Logit Bandits with General Value Functions
por: Zhang, Mengxiao, et al.
Publicado: (2024) -
Contextual Linear Bandits with Delay as Payoff
por: Zhang, Mengxiao, et al.
Publicado: (2025) -
Near-Optimal Last-Iterate Convergence for Zero-Sum Games with Bandit Feedback and Opponent Actions
por: Hait, Soumita, et al.
Publicado: (2026)