Off-Policy Learning in Large Action Spaces: Optimization Matters More Than Estimation
Fuente:
arXiv
Saved in:
| Main Authors: | Aouali, Imad, Sakhi, Otmane |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and Learning
by: Sakhi, Otmane, et al.
Published: (2024)
by: Sakhi, Otmane, et al.
Published: (2024)
Offline Contextual Bandit with Counterfactual Sample Identification
by: Gilotte, Alexandre, et al.
Published: (2025)
by: Gilotte, Alexandre, et al.
Published: (2025)
Sequential Off-Policy Learning with Logarithmic Smoothing
by: Haddouche, Maxime, et al.
Published: (2025)
by: Haddouche, Maxime, et al.
Published: (2025)
Bayesian Off-Policy Evaluation and Learning for Large Action Spaces
by: Aouali, Imad, et al.
Published: (2024)
by: Aouali, Imad, et al.
Published: (2024)
Exploiting Similarities in A/B Testing with Off-Policy Estimation
by: Sakhi, Otmane, et al.
Published: (2025)
by: Sakhi, Otmane, et al.
Published: (2025)
RecoAtlas: From Semantic Plausibility to Set-Level Utility in LLM Recommendation Agents
by: Aouali, Imad, et al.
Published: (2026)
by: Aouali, Imad, et al.
Published: (2026)
Probabilistic Rank and Reward: A Scalable Model for Slate Recommendation
by: Aouali, Imad, et al.
Published: (2022)
by: Aouali, Imad, et al.
Published: (2022)
Non-Linear Counterfactual Aggregate Optimization
by: Heymann, Benjamin, et al.
Published: (2025)
by: Heymann, Benjamin, et al.
Published: (2025)
Fast Slate Policy Optimization: Going Beyond Plackett-Luce
by: Sakhi, Otmane, et al.
Published: (2023)
by: Sakhi, Otmane, et al.
Published: (2023)
Learning to Bid in Repeated Second-Price Auctions with Dynamic Values and Aggregated Feedback
by: Heymann, Benjamin, et al.
Published: (2026)
by: Heymann, Benjamin, et al.
Published: (2026)
Diffusion Models Meet Contextual Bandits
by: Aouali, Imad
Published: (2024)
by: Aouali, Imad
Published: (2024)
RoiRL: Efficient, Self-Supervised Reasoning with Offline Iterative Reinforcement Learning
by: Arzhantsev, Aleksei, et al.
Published: (2025)
by: Arzhantsev, Aleksei, et al.
Published: (2025)
Self-Consistency via Marginal Sharpening
by: Arzhantsev, Aleksei, et al.
Published: (2026)
by: Arzhantsev, Aleksei, et al.
Published: (2026)
Unified PAC-Bayesian Study of Pessimism for Offline Policy Learning with Regularized Importance Sampling
by: Aouali, Imad, et al.
Published: (2024)
by: Aouali, Imad, et al.
Published: (2024)
Data Valuation for LLM Fine-Tuning: Efficient Shapley Value Approximation via Language Model Arithmetic
by: Tamine, Mélissa, et al.
Published: (2025)
by: Tamine, Mélissa, et al.
Published: (2025)
Prior-Dependent Allocations for Bayesian Fixed-Budget Best-Arm Identification in Structured Bandits
by: Nguyen, Nicolas, et al.
Published: (2024)
by: Nguyen, Nicolas, et al.
Published: (2024)
POTEC: Off-Policy Learning for Large Action Spaces via Two-Stage Policy Decomposition
by: Saito, Yuta, et al.
Published: (2024)
by: Saito, Yuta, et al.
Published: (2024)
Efficient Off-Policy Learning for High-Dimensional Action Spaces
by: Otto, Fabian, et al.
Published: (2024)
by: Otto, Fabian, et al.
Published: (2024)
Learning Action Embeddings for Off-Policy Evaluation
by: Cief, Matej, et al.
Published: (2023)
by: Cief, Matej, et al.
Published: (2023)
Context-Action Embedding Learning for Off-Policy Evaluation in Contextual Bandits
by: Chandak, Kushagra, et al.
Published: (2025)
by: Chandak, Kushagra, et al.
Published: (2025)
Behaviour Policy Optimization: Provably Lower Variance Return Estimates for Off-Policy Reinforcement Learning
by: Goodall, Alexander W., et al.
Published: (2025)
by: Goodall, Alexander W., et al.
Published: (2025)
More Than One Teacher: Adaptive Multi-Guidance Policy Optimization for Diverse Exploration
by: Yuan, Xiaoyang, et al.
Published: (2025)
by: Yuan, Xiaoyang, et al.
Published: (2025)
Log-Sum-Exponential Estimator for Off-Policy Evaluation and Learning
by: Behnamnia, Armin, et al.
Published: (2025)
by: Behnamnia, Armin, et al.
Published: (2025)
Off-OAB: Off-Policy Policy Gradient Method with Optimal Action-Dependent Baseline
by: Meng, Wenjia, et al.
Published: (2024)
by: Meng, Wenjia, et al.
Published: (2024)
RePO: Bridging On-Policy Learning and Off-Policy Knowledge through Rephrasing Policy Optimization
by: Xia, Linxuan, et al.
Published: (2026)
by: Xia, Linxuan, et al.
Published: (2026)
Doubly-Robust Off-Policy Evaluation with Estimated Logging Policy
by: Lee, Kyungbok, et al.
Published: (2024)
by: Lee, Kyungbok, et al.
Published: (2024)
Meta Off-Policy Estimation
by: Jeunen, Olivier
Published: (2025)
by: Jeunen, Olivier
Published: (2025)
Automated Off-Policy Estimator Selection via Supervised Learning
by: Felicioni, Nicolò, et al.
Published: (2024)
by: Felicioni, Nicolò, et al.
Published: (2024)
Pessimistic Off-Policy Optimization for Learning to Rank
by: Cief, Matej, et al.
Published: (2022)
by: Cief, Matej, et al.
Published: (2022)
Transductive Off-policy Proximal Policy Optimization
by: Gan, Yaozhong, et al.
Published: (2024)
by: Gan, Yaozhong, et al.
Published: (2024)
Coherent Off-Policy Improvement of Large Behavior Models with Learned Rewards
by: Scherer, Christian, et al.
Published: (2026)
by: Scherer, Christian, et al.
Published: (2026)
An Advantage-based Optimization Method for Reinforcement Learning in Large Action Space
by: Lin, Hai, et al.
Published: (2024)
by: Lin, Hai, et al.
Published: (2024)
Revisiting Group Relative Policy Optimization: Insights into On-Policy and Off-Policy Training
by: Mroueh, Youssef, et al.
Published: (2025)
by: Mroueh, Youssef, et al.
Published: (2025)
Labels Matter More Than Models: Rethinking the Unsupervised Paradigm in Time Series Anomaly Detection
by: Zhong, Zhijie, et al.
Published: (2025)
by: Zhong, Zhijie, et al.
Published: (2025)
Hyperparameter Optimization Can Even be Harmful in Off-Policy Learning and How to Deal with It
by: Saito, Yuta, et al.
Published: (2024)
by: Saito, Yuta, et al.
Published: (2024)
Off-Policy Evaluation of Slate Bandit Policies via Optimizing Abstraction
by: Kiyohara, Haruka, et al.
Published: (2024)
by: Kiyohara, Haruka, et al.
Published: (2024)
Off-Policy Learning with Limited Supply
by: Tanaka, Koichi, et al.
Published: (2026)
by: Tanaka, Koichi, et al.
Published: (2026)
More Than Memory Savings: Zeroth-Order Optimization Mitigates Forgetting in Continual Learning
by: Yu, Wanhao, et al.
Published: (2025)
by: Yu, Wanhao, et al.
Published: (2025)
Off-Policy Maximum Entropy RL with Future State and Action Visitation Measures
by: Bolland, Adrien, et al.
Published: (2024)
by: Bolland, Adrien, et al.
Published: (2024)
Off-Policy Evaluation of Ranking Policies via Embedding-Space User Behavior Modeling
by: Takahashi, Tatsuki, et al.
Published: (2025)
by: Takahashi, Tatsuki, et al.
Published: (2025)
Similar Items
-
Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and Learning
by: Sakhi, Otmane, et al.
Published: (2024) -
Offline Contextual Bandit with Counterfactual Sample Identification
by: Gilotte, Alexandre, et al.
Published: (2025) -
Sequential Off-Policy Learning with Logarithmic Smoothing
by: Haddouche, Maxime, et al.
Published: (2025) -
Bayesian Off-Policy Evaluation and Learning for Large Action Spaces
by: Aouali, Imad, et al.
Published: (2024) -
Exploiting Similarities in A/B Testing with Off-Policy Estimation
by: Sakhi, Otmane, et al.
Published: (2025)