Long-term Off-Policy Evaluation and Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Saito, Yuta, Abdollahpouri, Himan, Anderton, Jesse, Carterette, Ben, Lalmas, Mounia |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Calibrated Recommendations with Contextual Bandits
by: Feijer, Diego, et al.
Published: (2025)
by: Feijer, Diego, et al.
Published: (2025)
Off-Policy Evaluation and Learning for Matching Markets
by: Hayashi, Yudai, et al.
Published: (2025)
by: Hayashi, Yudai, et al.
Published: (2025)
Off-Policy Evaluation of Slate Bandit Policies via Optimizing Abstraction
by: Kiyohara, Haruka, et al.
Published: (2024)
by: Kiyohara, Haruka, et al.
Published: (2024)
Off-Policy Evaluation and Learning for Survival Outcomes under Censoring
by: Kubota, Kohsuke, et al.
Published: (2026)
by: Kubota, Kohsuke, et al.
Published: (2026)
Hyperparameter Optimization Can Even be Harmful in Off-Policy Learning and How to Deal with It
by: Saito, Yuta, et al.
Published: (2024)
by: Saito, Yuta, et al.
Published: (2024)
POTEC: Off-Policy Learning for Large Action Spaces via Two-Stage Policy Decomposition
by: Saito, Yuta, et al.
Published: (2024)
by: Saito, Yuta, et al.
Published: (2024)
Effective Off-Policy Evaluation and Learning in Contextual Combinatorial Bandits
by: Shimizu, Tatsuhiro, et al.
Published: (2024)
by: Shimizu, Tatsuhiro, et al.
Published: (2024)
Off-Policy Evaluation and Learning for the Future under Non-Stationarity
by: Shimizu, Tatsuhiro, et al.
Published: (2025)
by: Shimizu, Tatsuhiro, et al.
Published: (2025)
Off-Policy Learning with Limited Supply
by: Tanaka, Koichi, et al.
Published: (2026)
by: Tanaka, Koichi, et al.
Published: (2026)
Off-Policy Evaluation for Ranking Policies under Deterministic Logging Policies
by: Tanaka, Koichi, et al.
Published: (2026)
by: Tanaka, Koichi, et al.
Published: (2026)
SCOPE-RL: A Python Library for Offline Reinforcement Learning and Off-Policy Evaluation
by: Kiyohara, Haruka, et al.
Published: (2023)
by: Kiyohara, Haruka, et al.
Published: (2023)
Generalized User Representations for Transfer Learning
by: Fazelnia, Ghazal, et al.
Published: (2024)
by: Fazelnia, Ghazal, et al.
Published: (2024)
Towards Assessing and Benchmarking Risk-Return Tradeoff of Off-Policy Evaluation
by: Kiyohara, Haruka, et al.
Published: (2023)
by: Kiyohara, Haruka, et al.
Published: (2023)
Prompt-to-Slate: Diffusion Models for Prompt-Conditioned Slate Generation
by: Tomasi, Federico, et al.
Published: (2024)
by: Tomasi, Federico, et al.
Published: (2024)
A General Framework for Off-Policy Learning with Partially-Observed Reward
by: Takehi, Rikiya, et al.
Published: (2025)
by: Takehi, Rikiya, et al.
Published: (2025)
Towards Graph Foundation Models for Personalization
by: Damianou, Andreas, et al.
Published: (2024)
by: Damianou, Andreas, et al.
Published: (2024)
More is Less? A Simulation-Based Approach to Dynamic Interactions between Biases in Multimodal Models
by: Drissi, Mounia
Published: (2024)
by: Drissi, Mounia
Published: (2024)
MultiScale Contextual Bandits for Long Term Objectives
by: Rastogi, Richa, et al.
Published: (2025)
by: Rastogi, Richa, et al.
Published: (2025)
Learning Action Embeddings for Off-Policy Evaluation
by: Cief, Matej, et al.
Published: (2023)
by: Cief, Matej, et al.
Published: (2023)
Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and Learning
by: Sakhi, Otmane, et al.
Published: (2024)
by: Sakhi, Otmane, et al.
Published: (2024)
Log-Sum-Exponential Estimator for Off-Policy Evaluation and Learning
by: Behnamnia, Armin, et al.
Published: (2025)
by: Behnamnia, Armin, et al.
Published: (2025)
Cross-Validated Off-Policy Evaluation
by: Cief, Matej, et al.
Published: (2024)
by: Cief, Matej, et al.
Published: (2024)
Context-Action Embedding Learning for Off-Policy Evaluation in Contextual Bandits
by: Chandak, Kushagra, et al.
Published: (2025)
by: Chandak, Kushagra, et al.
Published: (2025)
DOLCE: Decomposing Off-Policy Evaluation/Learning into Lagged and Current Effects
by: Tamano, Shu
Published: (2025)
by: Tamano, Shu
Published: (2025)
Doubly-Robust Off-Policy Evaluation with Estimated Logging Policy
by: Lee, Kyungbok, et al.
Published: (2024)
by: Lee, Kyungbok, et al.
Published: (2024)
Evaluating Podcast Recommendations with Profile-Aware LLM-as-a-Judge
by: Fabbri, Francesco, et al.
Published: (2025)
by: Fabbri, Francesco, et al.
Published: (2025)
Kernel Metric Learning for In-Sample Off-Policy Evaluation of Deterministic RL Policies
by: Lee, Haanvid, et al.
Published: (2024)
by: Lee, Haanvid, et al.
Published: (2024)
Bayesian Off-Policy Evaluation and Learning for Large Action Spaces
by: Aouali, Imad, et al.
Published: (2024)
by: Aouali, Imad, et al.
Published: (2024)
Concept-driven Off Policy Evaluation
by: Majumdar, Ritam, et al.
Published: (2024)
by: Majumdar, Ritam, et al.
Published: (2024)
Clustering Context in Off-Policy Evaluation
by: Guzman-Olivares, Daniel, et al.
Published: (2025)
by: Guzman-Olivares, Daniel, et al.
Published: (2025)
Off-Policy Evaluation from Logged Human Feedback
by: Bhargava, Aniruddha, et al.
Published: (2024)
by: Bhargava, Aniruddha, et al.
Published: (2024)
Off-Policy Evaluation for Recommendations with Missing-Not-At-Random Rewards
by: Takahashi, Tatsuki, et al.
Published: (2025)
by: Takahashi, Tatsuki, et al.
Published: (2025)
Off-Policy Evaluation Under Nonignorable Missing Data
by: Wang, Han, et al.
Published: (2025)
by: Wang, Han, et al.
Published: (2025)
A Causal Lens for Learning Long-term Fair Policies
by: Lear, Jacob, et al.
Published: (2025)
by: Lear, Jacob, et al.
Published: (2025)
Pareto-Optimal Estimation and Policy Learning on Short-term and Long-term Treatment Effects
by: Wang, Yingrong, et al.
Published: (2024)
by: Wang, Yingrong, et al.
Published: (2024)
Formulating Reinforcement Learning for Human-Robot Collaboration through Off-Policy Evaluation
by: Singh, Saurav, et al.
Published: (2026)
by: Singh, Saurav, et al.
Published: (2026)
Sequential Off-Policy Learning with Logarithmic Smoothing
by: Haddouche, Maxime, et al.
Published: (2025)
by: Haddouche, Maxime, et al.
Published: (2025)
On the Reuse Bias in Off-Policy Reinforcement Learning
by: Ying, Chengyang, et al.
Published: (2022)
by: Ying, Chengyang, et al.
Published: (2022)
Efficient Dataset Selection for Continual Adaptation of Generative Recommenders
by: Jiao, Cathy, et al.
Published: (2026)
by: Jiao, Cathy, et al.
Published: (2026)
CANDOR: Counterfactual ANnotated DOubly Robust Off-Policy Evaluation
by: Mandyam, Aishwarya, et al.
Published: (2024)
by: Mandyam, Aishwarya, et al.
Published: (2024)
Similar Items
-
Calibrated Recommendations with Contextual Bandits
by: Feijer, Diego, et al.
Published: (2025) -
Off-Policy Evaluation and Learning for Matching Markets
by: Hayashi, Yudai, et al.
Published: (2025) -
Off-Policy Evaluation of Slate Bandit Policies via Optimizing Abstraction
by: Kiyohara, Haruka, et al.
Published: (2024) -
Off-Policy Evaluation and Learning for Survival Outcomes under Censoring
by: Kubota, Kohsuke, et al.
Published: (2026) -
Hyperparameter Optimization Can Even be Harmful in Off-Policy Learning and How to Deal with It
by: Saito, Yuta, et al.
Published: (2024)