SPEED: Experimental Design for Policy Evaluation in Linear Heteroscedastic Bandits
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mukherjee, Subhojyoti, Xie, Qiaomin, Hanna, Josiah, Nowak, Robert |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Pretraining Decision Transformers with Reward Prediction for In-Context Multi-task Structured Bandit Learning
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2024)
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2024)
SaVeR: Optimal Data Collection Strategy for Safe Policy Evaluation in Tabular MDP
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2024)
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2024)
Efficient and Interpretable Bandit Algorithms
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2023)
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2023)
Stable Offline Value Function Learning with Bisimulation-based Representations
von: Pavse, Brahma S., et al.
Veröffentlicht: (2024)
von: Pavse, Brahma S., et al.
Veröffentlicht: (2024)
Learning to Stabilize Online Reinforcement Learning in Unbounded State Spaces
von: Pavse, Brahma S., et al.
Veröffentlicht: (2023)
von: Pavse, Brahma S., et al.
Veröffentlicht: (2023)
Off-Policy Evaluation from Logged Human Feedback
von: Bhargava, Aniruddha, et al.
Veröffentlicht: (2024)
von: Bhargava, Aniruddha, et al.
Veröffentlicht: (2024)
On-Policy Policy Gradient Reinforcement Learning Without On-Policy Sampling
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2023)
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2023)
An Empirical Study on the Power of Future Prediction in Partially Observable Environments
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024)
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024)
Experimental Design for Active Transductive Inference in Large Language Models
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2024)
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2024)
Bias and Extrapolation in Markovian Linear Stochastic Approximation with Constant Stepsizes
von: Huo, Dongyan, et al.
Veröffentlicht: (2022)
von: Huo, Dongyan, et al.
Veröffentlicht: (2022)
Partial Policy Gradients for RL in LLMs
von: Mathur, Puneet, et al.
Veröffentlicht: (2026)
von: Mathur, Puneet, et al.
Veröffentlicht: (2026)
Centralized Adaptive Sampling for Reliable Co-Training of Independent Multi-Agent Policies
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2025)
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2025)
Wasserstein-p Central Limit Theorem Rates: From Local Dependence to Markov Chains
von: Zhang, Yixuan, et al.
Veröffentlicht: (2026)
von: Zhang, Yixuan, et al.
Veröffentlicht: (2026)
Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation
von: Zhou, Hongyi, et al.
Veröffentlicht: (2025)
von: Zhou, Hongyi, et al.
Veröffentlicht: (2025)
Robust Causal Bandits for Linear Models
von: Yan, Zirui, et al.
Veröffentlicht: (2023)
von: Yan, Zirui, et al.
Veröffentlicht: (2023)
A Reduction Algorithm for Markovian Contextual Linear Bandits
von: Buyukkalayci, Kaan, et al.
Veröffentlicht: (2026)
von: Buyukkalayci, Kaan, et al.
Veröffentlicht: (2026)
Constant Stepsize Q-learning: Distributional Convergence, Bias and Extrapolation
von: Zhang, Yixuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yixuan, et al.
Veröffentlicht: (2024)
Coupling-based Convergence Diagnostic and Stepsize Scheme for Stochastic Gradient Descent
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
Restless Bandits with Average Reward: Breaking the Uniform Global Attractor Assumption
von: Hong, Yige, et al.
Veröffentlicht: (2023)
von: Hong, Yige, et al.
Veröffentlicht: (2023)
Experimental Design for Semiparametric Bandits
von: Kim, Seok-Jin, et al.
Veröffentlicht: (2025)
von: Kim, Seok-Jin, et al.
Veröffentlicht: (2025)
Improved Bound for Robust Causal Bandits with Linear Models
von: Yan, Zirui, et al.
Veröffentlicht: (2024)
von: Yan, Zirui, et al.
Veröffentlicht: (2024)
Achieving Exponential Asymptotic Optimality in Average-Reward Restless Bandits without Global Attractor Assumption
von: Hong, Yige, et al.
Veröffentlicht: (2024)
von: Hong, Yige, et al.
Veröffentlicht: (2024)
Two-Timescale Linear Stochastic Approximation: Constant Stepsizes Go a Long Way
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024)
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024)
Contextual Online Pricing with (Biased) Offline Data
von: Zhang, Yixuan, et al.
Veröffentlicht: (2025)
von: Zhang, Yixuan, et al.
Veröffentlicht: (2025)
On the Peril of (Even a Little) Nonstationarity in Satisficing Regret Minimization
von: Zhang, Yixuan, et al.
Veröffentlicht: (2026)
von: Zhang, Yixuan, et al.
Veröffentlicht: (2026)
Optimal Design for Human Preference Elicitation
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2024)
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2024)
Adaptive Exploration for Data-Efficient General Value Function Evaluations
von: Jain, Arushi, et al.
Veröffentlicht: (2024)
von: Jain, Arushi, et al.
Veröffentlicht: (2024)
Agentic Planning with Reasoning for Image Styling via Offline RL
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2026)
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2026)
Understanding when Dynamics-Invariant Data Augmentations Benefit Model-Free Reinforcement Learning Updates
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2023)
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2023)
Multi-Objective Alignment of Large Language Models Through Hypervolume Maximization
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2024)
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2024)
Logits are All We Need to Adapt Closed Models
von: Hiranandani, Gaurush, et al.
Veröffentlicht: (2025)
von: Hiranandani, Gaurush, et al.
Veröffentlicht: (2025)
Unichain and Aperiodicity are Sufficient for Asymptotic Optimality of Average-Reward Restless Bandits
von: Hong, Yige, et al.
Veröffentlicht: (2024)
von: Hong, Yige, et al.
Veröffentlicht: (2024)
Off-Policy Evaluation of Slate Bandit Policies via Optimizing Abstraction
von: Kiyohara, Haruka, et al.
Veröffentlicht: (2024)
von: Kiyohara, Haruka, et al.
Veröffentlicht: (2024)
Scaling Federated Linear Contextual Bandits via Sketching
von: Yang, Hantao, et al.
Veröffentlicht: (2026)
von: Yang, Hantao, et al.
Veröffentlicht: (2026)
New Classes of the Greedy-Applicable Arm Feature Distributions in the Sparse Linear Bandit Problem
von: Ichikawa, Koji, et al.
Veröffentlicht: (2023)
von: Ichikawa, Koji, et al.
Veröffentlicht: (2023)
Bandit Max-Min Fair Allocation
von: Harada, Tsubasa, et al.
Veröffentlicht: (2025)
von: Harada, Tsubasa, et al.
Veröffentlicht: (2025)
C-kNN-LSH: A Nearest-Neighbor Algorithm for Sequential Counterfactual Inference
von: Wang, Jing, et al.
Veröffentlicht: (2026)
von: Wang, Jing, et al.
Veröffentlicht: (2026)
Wasserstein Distributionally Robust Policy Evaluation and Learning for Contextual Bandits
von: Shen, Yi, et al.
Veröffentlicht: (2023)
von: Shen, Yi, et al.
Veröffentlicht: (2023)
Design Experiments to Compare Multi-armed Bandit Algorithms
von: Meng, Huiling, et al.
Veröffentlicht: (2026)
von: Meng, Huiling, et al.
Veröffentlicht: (2026)
Distributionally Robust Multi-Task Reinforcement Learning via Adaptive Task Sampling
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2026)
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Pretraining Decision Transformers with Reward Prediction for In-Context Multi-task Structured Bandit Learning
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2024) -
SaVeR: Optimal Data Collection Strategy for Safe Policy Evaluation in Tabular MDP
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2024) -
Efficient and Interpretable Bandit Algorithms
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2023) -
Stable Offline Value Function Learning with Bisimulation-based Representations
von: Pavse, Brahma S., et al.
Veröffentlicht: (2024) -
Learning to Stabilize Online Reinforcement Learning in Unbounded State Spaces
von: Pavse, Brahma S., et al.
Veröffentlicht: (2023)