Second Order Bounds for Contextual Bandits with Function Approximation
Fuente:
arXiv
Guardado en:
| Autor principal: | Pacchiano, Aldo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Experiment Planning with Function Approximation
por: Pacchiano, Aldo, et al.
Publicado: (2024)
por: Pacchiano, Aldo, et al.
Publicado: (2024)
Learning Rate-Free Reinforcement Learning: A Case for Model Selection with Non-Stationary Objectives
por: Afshar, Aida, et al.
Publicado: (2024)
por: Afshar, Aida, et al.
Publicado: (2024)
Adaptive Exploration for Multi-Reward Multi-Policy Evaluation
por: Russo, Alessio, et al.
Publicado: (2025)
por: Russo, Alessio, et al.
Publicado: (2025)
Improved Training Mechanism for Reinforcement Learning via Online Model Selection
por: Afshar, Aida, et al.
Publicado: (2025)
por: Afshar, Aida, et al.
Publicado: (2025)
Contextual Bandits with Stage-wise Constraints
por: Pacchiano, Aldo, et al.
Publicado: (2024)
por: Pacchiano, Aldo, et al.
Publicado: (2024)
State-free Reinforcement Learning
por: Chen, Mingyu, et al.
Publicado: (2024)
por: Chen, Mingyu, et al.
Publicado: (2024)
Data-Driven Online Model Selection With Regret Guarantees
por: Pacchiano, Aldo, et al.
Publicado: (2023)
por: Pacchiano, Aldo, et al.
Publicado: (2023)
In-Context Learning for Pure Exploration
por: Russo, Alessio, et al.
Publicado: (2025)
por: Russo, Alessio, et al.
Publicado: (2025)
Variance-Dependent Regret Lower Bounds for Contextual Bandits
por: He, Jiafan, et al.
Publicado: (2025)
por: He, Jiafan, et al.
Publicado: (2025)
Multiple-policy Evaluation via Density Estimation
por: Chen, Yilei, et al.
Publicado: (2024)
por: Chen, Yilei, et al.
Publicado: (2024)
A Theoretical Framework for Partially Observed Reward-States in RLHF
por: Kausik, Chinmaya, et al.
Publicado: (2024)
por: Kausik, Chinmaya, et al.
Publicado: (2024)
In-Context Learning for Pure Exploration in Continuous Spaces
por: Russo, Alessio, et al.
Publicado: (2026)
por: Russo, Alessio, et al.
Publicado: (2026)
Provable Interactive Learning with Hindsight Instruction Feedback
por: Misra, Dipendra, et al.
Publicado: (2024)
por: Misra, Dipendra, et al.
Publicado: (2024)
The Good, the Bad, and the Sampled: a No-Regret Approach to Safe Online Classification
por: Baharav, Tavor Z., et al.
Publicado: (2025)
por: Baharav, Tavor Z., et al.
Publicado: (2025)
Tree Ensembles for Contextual Bandits
por: Nilsson, Hannes, et al.
Publicado: (2024)
por: Nilsson, Hannes, et al.
Publicado: (2024)
Causal Contextual Bandits with Adaptive Context
por: Madhavan, Rahul, et al.
Publicado: (2024)
por: Madhavan, Rahul, et al.
Publicado: (2024)
Diffusion Models Meet Contextual Bandits
por: Aouali, Imad
Publicado: (2024)
por: Aouali, Imad
Publicado: (2024)
Learning When to Trust in Contextual Bandits
por: Ghasemi, Majid, et al.
Publicado: (2026)
por: Ghasemi, Majid, et al.
Publicado: (2026)
Meet Me at the Arm: The Cooperative Multi-Armed Bandits Problem with Shareable Arms
por: Hu, Xinyi, et al.
Publicado: (2025)
por: Hu, Xinyi, et al.
Publicado: (2025)
A Contextual Combinatorial Bandit Approach to Negotiation
por: Li, Yexin, et al.
Publicado: (2024)
por: Li, Yexin, et al.
Publicado: (2024)
Federated Linear Contextual Bandits with Heterogeneous Clients
por: Blaser, Ethan, et al.
Publicado: (2024)
por: Blaser, Ethan, et al.
Publicado: (2024)
Conservative Contextual Bandits: Beyond Linear Representations
por: Deb, Rohan, et al.
Publicado: (2024)
por: Deb, Rohan, et al.
Publicado: (2024)
Linear Contextual Bandits with Hybrid Payoff: Revisited
por: Das, Nirjhar, et al.
Publicado: (2024)
por: Das, Nirjhar, et al.
Publicado: (2024)
The Sample Complexity of Multiclass and Sparse Contextual Bandits
por: Erez, Liad, et al.
Publicado: (2026)
por: Erez, Liad, et al.
Publicado: (2026)
Contextual Combinatorial Bandits with Probabilistically Triggered Arms
por: Liu, Xutong, et al.
Publicado: (2023)
por: Liu, Xutong, et al.
Publicado: (2023)
Active Preference Optimization for Sample Efficient RLHF
por: Das, Nirjhar, et al.
Publicado: (2024)
por: Das, Nirjhar, et al.
Publicado: (2024)
Principled Fine-tuning of LLMs from User-Edits: A Medley of Preference, Supervision, and Reward
por: Misra, Dipendra, et al.
Publicado: (2026)
por: Misra, Dipendra, et al.
Publicado: (2026)
Leveraging Offline Data in Linear Latent Contextual Bandits
por: Kausik, Chinmaya, et al.
Publicado: (2024)
por: Kausik, Chinmaya, et al.
Publicado: (2024)
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
por: Lu, Xiaodong, et al.
Publicado: (2026)
por: Lu, Xiaodong, et al.
Publicado: (2026)
ORSO: Accelerating Reward Design via Online Reward Selection and Policy Optimization
por: Zhang, Chen Bo Calvin, et al.
Publicado: (2024)
por: Zhang, Chen Bo Calvin, et al.
Publicado: (2024)
When Less is Enough: Efficient Inference via Collaborative Reasoning
por: Chen, Yilei, et al.
Publicado: (2026)
por: Chen, Yilei, et al.
Publicado: (2026)
From Contextual Combinatorial Semi-Bandits to Bandit List Classification: Improved Sample Complexity with Sparse Rewards
por: Erez, Liad, et al.
Publicado: (2025)
por: Erez, Liad, et al.
Publicado: (2025)
Effective Off-Policy Evaluation and Learning in Contextual Combinatorial Bandits
por: Shimizu, Tatsuhiro, et al.
Publicado: (2024)
por: Shimizu, Tatsuhiro, et al.
Publicado: (2024)
Federated Contextual Cascading Bandits with Asynchronous Communication and Heterogeneous Users
por: Yang, Hantao, et al.
Publicado: (2024)
por: Yang, Hantao, et al.
Publicado: (2024)
Provable Anytime Ensemble Sampling Algorithms in Nonlinear Contextual Bandits
por: Sun, Jiazheng, et al.
Publicado: (2025)
por: Sun, Jiazheng, et al.
Publicado: (2025)
COBRA: Contextual Bandit Algorithm for Ensuring Truthful Strategic Agents
por: Verma, Arun, et al.
Publicado: (2025)
por: Verma, Arun, et al.
Publicado: (2025)
Calibration-Gated LLM Pseudo-Observations for Online Contextual Bandits
por: Pershin, Maksim, et al.
Publicado: (2026)
por: Pershin, Maksim, et al.
Publicado: (2026)
Misspecified $Q$-Learning with Sparse Linear Function Approximation: Tight Bounds on Approximation Error
por: Du, Ally Yalei, et al.
Publicado: (2024)
por: Du, Ally Yalei, et al.
Publicado: (2024)
Contextual Linear Bandits under Noisy Features: Towards Bayesian Oracles
por: Kim, Jung-hun, et al.
Publicado: (2017)
por: Kim, Jung-hun, et al.
Publicado: (2017)
On the Hardness of Bandit Learning
por: Brukhim, Nataly, et al.
Publicado: (2025)
por: Brukhim, Nataly, et al.
Publicado: (2025)
Ejemplares similares
-
Experiment Planning with Function Approximation
por: Pacchiano, Aldo, et al.
Publicado: (2024) -
Learning Rate-Free Reinforcement Learning: A Case for Model Selection with Non-Stationary Objectives
por: Afshar, Aida, et al.
Publicado: (2024) -
Adaptive Exploration for Multi-Reward Multi-Policy Evaluation
por: Russo, Alessio, et al.
Publicado: (2025) -
Improved Training Mechanism for Reinforcement Learning via Online Model Selection
por: Afshar, Aida, et al.
Publicado: (2025) -
Contextual Bandits with Stage-wise Constraints
por: Pacchiano, Aldo, et al.
Publicado: (2024)