Context-Action Embedding Learning for Off-Policy Evaluation in Contextual Bandits
Fuente:
arXiv
Saved in:
| Main Authors: | Chandak, Kushagra, Liu, Vincent, Lee, Haanvid |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Kernel Metric Learning for In-Sample Off-Policy Evaluation of Deterministic RL Policies
by: Lee, Haanvid, et al.
Published: (2024)
by: Lee, Haanvid, et al.
Published: (2024)
Effective Off-Policy Evaluation and Learning in Contextual Combinatorial Bandits
by: Shimizu, Tatsuhiro, et al.
Published: (2024)
by: Shimizu, Tatsuhiro, et al.
Published: (2024)
PAC Off-Policy Prediction of Contextual Bandits
by: Wan, Yilong, et al.
Published: (2025)
by: Wan, Yilong, et al.
Published: (2025)
Learning Action Embeddings for Off-Policy Evaluation
by: Cief, Matej, et al.
Published: (2023)
by: Cief, Matej, et al.
Published: (2023)
Short-Long Policy Evaluation with Novel Actions
by: Nam, Hyunji Alex, et al.
Published: (2024)
by: Nam, Hyunji Alex, et al.
Published: (2024)
Optimal Baseline Corrections for Off-Policy Contextual Bandits
by: Gupta, Shashank, et al.
Published: (2024)
by: Gupta, Shashank, et al.
Published: (2024)
Off-Policy Evaluation via Adaptive Weighting with Data from Contextual Bandits
by: Zhan, Ruohan, et al.
Published: (2021)
by: Zhan, Ruohan, et al.
Published: (2021)
Wasserstein Distributionally Robust Policy Evaluation and Learning for Contextual Bandits
by: Shen, Yi, et al.
Published: (2023)
by: Shen, Yi, et al.
Published: (2023)
When is Off-Policy Evaluation (Reward Modeling) Useful in Contextual Bandits? A Data-Centric Perspective
by: Sun, Hao, et al.
Published: (2023)
by: Sun, Hao, et al.
Published: (2023)
Off-Policy Evaluation of Slate Bandit Policies via Optimizing Abstraction
by: Kiyohara, Haruka, et al.
Published: (2024)
by: Kiyohara, Haruka, et al.
Published: (2024)
Distributionally Robust Policy Evaluation under General Covariate Shift in Contextual Bandits
by: Guo, Yihong, et al.
Published: (2024)
by: Guo, Yihong, et al.
Published: (2024)
Pessimistic Risk-Aware Policy Learning in Contextual Bandits
by: Wan, Yilong, et al.
Published: (2026)
by: Wan, Yilong, et al.
Published: (2026)
Contextual Bandits for Unbounded Context Distributions
by: Zhao, Puning, et al.
Published: (2024)
by: Zhao, Puning, et al.
Published: (2024)
Constrained Contextual Bandits with Adversarial Contexts
by: Sarkar, Dhruv, et al.
Published: (2026)
by: Sarkar, Dhruv, et al.
Published: (2026)
Offline Contextual Bandits in the Presence of New Actions
by: Kishimoto, Ren, et al.
Published: (2026)
by: Kishimoto, Ren, et al.
Published: (2026)
Learning with Incomplete Context: Linear Contextual Bandits with Pretrained Imputation
by: Yan, Hao, et al.
Published: (2025)
by: Yan, Hao, et al.
Published: (2025)
Causal Contextual Bandits with Adaptive Context
by: Madhavan, Rahul, et al.
Published: (2024)
by: Madhavan, Rahul, et al.
Published: (2024)
Optimal Regret for Policy Optimization in Contextual Bandits
by: Levy, Orin, et al.
Published: (2026)
by: Levy, Orin, et al.
Published: (2026)
Clustering Context in Off-Policy Evaluation
by: Guzman-Olivares, Daniel, et al.
Published: (2025)
by: Guzman-Olivares, Daniel, et al.
Published: (2025)
Bayesian Off-Policy Evaluation and Learning for Large Action Spaces
by: Aouali, Imad, et al.
Published: (2024)
by: Aouali, Imad, et al.
Published: (2024)
Optimizing Warfarin Dosing Using Contextual Bandit: An Offline Policy Learning and Evaluation Method
by: Huang, Yong, et al.
Published: (2024)
by: Huang, Yong, et al.
Published: (2024)
Regret Minimization via Saddle Point Optimization
by: Kirschner, Johannes, et al.
Published: (2024)
by: Kirschner, Johannes, et al.
Published: (2024)
High Probability Bound for Cross-Learning Contextual Bandits with Unknown Context Distributions
by: Huang, Ruiyuan, et al.
Published: (2024)
by: Huang, Ruiyuan, et al.
Published: (2024)
Log-Sum-Exponential Estimator for Off-Policy Evaluation and Learning
by: Behnamnia, Armin, et al.
Published: (2025)
by: Behnamnia, Armin, et al.
Published: (2025)
Active Context Selection Improves Simple Regret in Contextual Bandits
by: Shahverdikondori, Mohammad, et al.
Published: (2026)
by: Shahverdikondori, Mohammad, et al.
Published: (2026)
Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments
by: Verma, Abhishek, et al.
Published: (2025)
by: Verma, Abhishek, et al.
Published: (2025)
Bayesian Inference of Contextual Bandit Policies via Empirical Likelihood
by: Ouyang, Jiangrong, et al.
Published: (2026)
by: Ouyang, Jiangrong, et al.
Published: (2026)
Off-Policy Evaluation of Ranking Policies via Embedding-Space User Behavior Modeling
by: Takahashi, Tatsuki, et al.
Published: (2025)
by: Takahashi, Tatsuki, et al.
Published: (2025)
Doubly-Robust Off-Policy Evaluation with Estimated Logging Policy
by: Lee, Kyungbok, et al.
Published: (2024)
by: Lee, Kyungbok, et al.
Published: (2024)
Efficient Off-Policy Learning for High-Dimensional Action Spaces
by: Otto, Fabian, et al.
Published: (2024)
by: Otto, Fabian, et al.
Published: (2024)
Off-Policy Evaluation Using Information Borrowing and Context-Based Switching
by: Dasgupta, Sutanoy, et al.
Published: (2021)
by: Dasgupta, Sutanoy, et al.
Published: (2021)
Online Multi-LLM Selection via Contextual Bandits under Unstructured Context Evolution
by: Poon, Manhin, et al.
Published: (2025)
by: Poon, Manhin, et al.
Published: (2025)
Active Learning for Stochastic Contextual Linear Bandits
by: Brunskill, Emma, et al.
Published: (2026)
by: Brunskill, Emma, et al.
Published: (2026)
Cramming Contextual Bandits for On-policy Statistical Evaluation
by: Jia, Zeyang, et al.
Published: (2024)
by: Jia, Zeyang, et al.
Published: (2024)
Long-term Off-Policy Evaluation and Learning
by: Saito, Yuta, et al.
Published: (2024)
by: Saito, Yuta, et al.
Published: (2024)
Queue Length Regret Bounds for Contextual Queueing Bandits
by: Bae, Seoungbin, et al.
Published: (2026)
by: Bae, Seoungbin, et al.
Published: (2026)
Group-Sensitive Offline Contextual Bandits
by: Guo, Yihong, et al.
Published: (2025)
by: Guo, Yihong, et al.
Published: (2025)
Learning When to Trust in Contextual Bandits
by: Ghasemi, Majid, et al.
Published: (2026)
by: Ghasemi, Majid, et al.
Published: (2026)
Quantum-Enhanced Neural Contextual Bandit Algorithms
by: Huang, Yuqi, et al.
Published: (2026)
by: Huang, Yuqi, et al.
Published: (2026)
Learning to Route and Schedule LLMs from User Retrials via Contextual Queueing Bandits
by: Bae, Seoungbin, et al.
Published: (2026)
by: Bae, Seoungbin, et al.
Published: (2026)
Similar Items
-
Kernel Metric Learning for In-Sample Off-Policy Evaluation of Deterministic RL Policies
by: Lee, Haanvid, et al.
Published: (2024) -
Effective Off-Policy Evaluation and Learning in Contextual Combinatorial Bandits
by: Shimizu, Tatsuhiro, et al.
Published: (2024) -
PAC Off-Policy Prediction of Contextual Bandits
by: Wan, Yilong, et al.
Published: (2025) -
Learning Action Embeddings for Off-Policy Evaluation
by: Cief, Matej, et al.
Published: (2023) -
Short-Long Policy Evaluation with Novel Actions
by: Nam, Hyunji Alex, et al.
Published: (2024)