SPEED: Experimental Design for Policy Evaluation in Linear Heteroscedastic Bandits
Fuente:
arXiv
Guardado en:
| Autores principales: | Mukherjee, Subhojyoti, Xie, Qiaomin, Hanna, Josiah, Nowak, Robert |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Pretraining Decision Transformers with Reward Prediction for In-Context Multi-task Structured Bandit Learning
por: Mukherjee, Subhojyoti, et al.
Publicado: (2024)
por: Mukherjee, Subhojyoti, et al.
Publicado: (2024)
SaVeR: Optimal Data Collection Strategy for Safe Policy Evaluation in Tabular MDP
por: Mukherjee, Subhojyoti, et al.
Publicado: (2024)
por: Mukherjee, Subhojyoti, et al.
Publicado: (2024)
Efficient and Interpretable Bandit Algorithms
por: Mukherjee, Subhojyoti, et al.
Publicado: (2023)
por: Mukherjee, Subhojyoti, et al.
Publicado: (2023)
Stable Offline Value Function Learning with Bisimulation-based Representations
por: Pavse, Brahma S., et al.
Publicado: (2024)
por: Pavse, Brahma S., et al.
Publicado: (2024)
Learning to Stabilize Online Reinforcement Learning in Unbounded State Spaces
por: Pavse, Brahma S., et al.
Publicado: (2023)
por: Pavse, Brahma S., et al.
Publicado: (2023)
Off-Policy Evaluation from Logged Human Feedback
por: Bhargava, Aniruddha, et al.
Publicado: (2024)
por: Bhargava, Aniruddha, et al.
Publicado: (2024)
On-Policy Policy Gradient Reinforcement Learning Without On-Policy Sampling
por: Corrado, Nicholas E., et al.
Publicado: (2023)
por: Corrado, Nicholas E., et al.
Publicado: (2023)
An Empirical Study on the Power of Future Prediction in Partially Observable Environments
por: Kwon, Jeongyeol, et al.
Publicado: (2024)
por: Kwon, Jeongyeol, et al.
Publicado: (2024)
Experimental Design for Active Transductive Inference in Large Language Models
por: Mukherjee, Subhojyoti, et al.
Publicado: (2024)
por: Mukherjee, Subhojyoti, et al.
Publicado: (2024)
Bias and Extrapolation in Markovian Linear Stochastic Approximation with Constant Stepsizes
por: Huo, Dongyan, et al.
Publicado: (2022)
por: Huo, Dongyan, et al.
Publicado: (2022)
Partial Policy Gradients for RL in LLMs
por: Mathur, Puneet, et al.
Publicado: (2026)
por: Mathur, Puneet, et al.
Publicado: (2026)
Centralized Adaptive Sampling for Reliable Co-Training of Independent Multi-Agent Policies
por: Corrado, Nicholas E., et al.
Publicado: (2025)
por: Corrado, Nicholas E., et al.
Publicado: (2025)
Wasserstein-p Central Limit Theorem Rates: From Local Dependence to Markov Chains
por: Zhang, Yixuan, et al.
Publicado: (2026)
por: Zhang, Yixuan, et al.
Publicado: (2026)
Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation
por: Zhou, Hongyi, et al.
Publicado: (2025)
por: Zhou, Hongyi, et al.
Publicado: (2025)
Robust Causal Bandits for Linear Models
por: Yan, Zirui, et al.
Publicado: (2023)
por: Yan, Zirui, et al.
Publicado: (2023)
A Reduction Algorithm for Markovian Contextual Linear Bandits
por: Buyukkalayci, Kaan, et al.
Publicado: (2026)
por: Buyukkalayci, Kaan, et al.
Publicado: (2026)
Constant Stepsize Q-learning: Distributional Convergence, Bias and Extrapolation
por: Zhang, Yixuan, et al.
Publicado: (2024)
por: Zhang, Yixuan, et al.
Publicado: (2024)
Coupling-based Convergence Diagnostic and Stepsize Scheme for Stochastic Gradient Descent
por: Li, Xiang, et al.
Publicado: (2024)
por: Li, Xiang, et al.
Publicado: (2024)
Restless Bandits with Average Reward: Breaking the Uniform Global Attractor Assumption
por: Hong, Yige, et al.
Publicado: (2023)
por: Hong, Yige, et al.
Publicado: (2023)
Experimental Design for Semiparametric Bandits
por: Kim, Seok-Jin, et al.
Publicado: (2025)
por: Kim, Seok-Jin, et al.
Publicado: (2025)
Improved Bound for Robust Causal Bandits with Linear Models
por: Yan, Zirui, et al.
Publicado: (2024)
por: Yan, Zirui, et al.
Publicado: (2024)
Achieving Exponential Asymptotic Optimality in Average-Reward Restless Bandits without Global Attractor Assumption
por: Hong, Yige, et al.
Publicado: (2024)
por: Hong, Yige, et al.
Publicado: (2024)
Two-Timescale Linear Stochastic Approximation: Constant Stepsizes Go a Long Way
por: Kwon, Jeongyeol, et al.
Publicado: (2024)
por: Kwon, Jeongyeol, et al.
Publicado: (2024)
Contextual Online Pricing with (Biased) Offline Data
por: Zhang, Yixuan, et al.
Publicado: (2025)
por: Zhang, Yixuan, et al.
Publicado: (2025)
On the Peril of (Even a Little) Nonstationarity in Satisficing Regret Minimization
por: Zhang, Yixuan, et al.
Publicado: (2026)
por: Zhang, Yixuan, et al.
Publicado: (2026)
Optimal Design for Human Preference Elicitation
por: Mukherjee, Subhojyoti, et al.
Publicado: (2024)
por: Mukherjee, Subhojyoti, et al.
Publicado: (2024)
Adaptive Exploration for Data-Efficient General Value Function Evaluations
por: Jain, Arushi, et al.
Publicado: (2024)
por: Jain, Arushi, et al.
Publicado: (2024)
Agentic Planning with Reasoning for Image Styling via Offline RL
por: Mukherjee, Subhojyoti, et al.
Publicado: (2026)
por: Mukherjee, Subhojyoti, et al.
Publicado: (2026)
Understanding when Dynamics-Invariant Data Augmentations Benefit Model-Free Reinforcement Learning Updates
por: Corrado, Nicholas E., et al.
Publicado: (2023)
por: Corrado, Nicholas E., et al.
Publicado: (2023)
Multi-Objective Alignment of Large Language Models Through Hypervolume Maximization
por: Mukherjee, Subhojyoti, et al.
Publicado: (2024)
por: Mukherjee, Subhojyoti, et al.
Publicado: (2024)
Logits are All We Need to Adapt Closed Models
por: Hiranandani, Gaurush, et al.
Publicado: (2025)
por: Hiranandani, Gaurush, et al.
Publicado: (2025)
Unichain and Aperiodicity are Sufficient for Asymptotic Optimality of Average-Reward Restless Bandits
por: Hong, Yige, et al.
Publicado: (2024)
por: Hong, Yige, et al.
Publicado: (2024)
Off-Policy Evaluation of Slate Bandit Policies via Optimizing Abstraction
por: Kiyohara, Haruka, et al.
Publicado: (2024)
por: Kiyohara, Haruka, et al.
Publicado: (2024)
Scaling Federated Linear Contextual Bandits via Sketching
por: Yang, Hantao, et al.
Publicado: (2026)
por: Yang, Hantao, et al.
Publicado: (2026)
New Classes of the Greedy-Applicable Arm Feature Distributions in the Sparse Linear Bandit Problem
por: Ichikawa, Koji, et al.
Publicado: (2023)
por: Ichikawa, Koji, et al.
Publicado: (2023)
Bandit Max-Min Fair Allocation
por: Harada, Tsubasa, et al.
Publicado: (2025)
por: Harada, Tsubasa, et al.
Publicado: (2025)
C-kNN-LSH: A Nearest-Neighbor Algorithm for Sequential Counterfactual Inference
por: Wang, Jing, et al.
Publicado: (2026)
por: Wang, Jing, et al.
Publicado: (2026)
Wasserstein Distributionally Robust Policy Evaluation and Learning for Contextual Bandits
por: Shen, Yi, et al.
Publicado: (2023)
por: Shen, Yi, et al.
Publicado: (2023)
Design Experiments to Compare Multi-armed Bandit Algorithms
por: Meng, Huiling, et al.
Publicado: (2026)
por: Meng, Huiling, et al.
Publicado: (2026)
Distributionally Robust Multi-Task Reinforcement Learning via Adaptive Task Sampling
por: Corrado, Nicholas E., et al.
Publicado: (2026)
por: Corrado, Nicholas E., et al.
Publicado: (2026)
Ejemplares similares
-
Pretraining Decision Transformers with Reward Prediction for In-Context Multi-task Structured Bandit Learning
por: Mukherjee, Subhojyoti, et al.
Publicado: (2024) -
SaVeR: Optimal Data Collection Strategy for Safe Policy Evaluation in Tabular MDP
por: Mukherjee, Subhojyoti, et al.
Publicado: (2024) -
Efficient and Interpretable Bandit Algorithms
por: Mukherjee, Subhojyoti, et al.
Publicado: (2023) -
Stable Offline Value Function Learning with Bisimulation-based Representations
por: Pavse, Brahma S., et al.
Publicado: (2024) -
Learning to Stabilize Online Reinforcement Learning in Unbounded State Spaces
por: Pavse, Brahma S., et al.
Publicado: (2023)