Offline Contextual Bandit with Counterfactual Sample Identification
Fuente:
arXiv
Saved in:
| Main Authors: | Gilotte, Alexandre, Sakhi, Otmane, Aouali, Imad, Heymann, Benjamin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RecoAtlas: From Semantic Plausibility to Set-Level Utility in LLM Recommendation Agents
by: Aouali, Imad, et al.
Published: (2026)
by: Aouali, Imad, et al.
Published: (2026)
Off-Policy Learning in Large Action Spaces: Optimization Matters More Than Estimation
by: Aouali, Imad, et al.
Published: (2025)
by: Aouali, Imad, et al.
Published: (2025)
Non-Linear Counterfactual Aggregate Optimization
by: Heymann, Benjamin, et al.
Published: (2025)
by: Heymann, Benjamin, et al.
Published: (2025)
Exploiting Similarities in A/B Testing with Off-Policy Estimation
by: Sakhi, Otmane, et al.
Published: (2025)
by: Sakhi, Otmane, et al.
Published: (2025)
Learning to Bid in Repeated Second-Price Auctions with Dynamic Values and Aggregated Feedback
by: Heymann, Benjamin, et al.
Published: (2026)
by: Heymann, Benjamin, et al.
Published: (2026)
Diffusion Models Meet Contextual Bandits
by: Aouali, Imad
Published: (2024)
by: Aouali, Imad
Published: (2024)
Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and Learning
by: Sakhi, Otmane, et al.
Published: (2024)
by: Sakhi, Otmane, et al.
Published: (2024)
Data Valuation for LLM Fine-Tuning: Efficient Shapley Value Approximation via Language Model Arithmetic
by: Tamine, Mélissa, et al.
Published: (2025)
by: Tamine, Mélissa, et al.
Published: (2025)
Probabilistic Rank and Reward: A Scalable Model for Slate Recommendation
by: Aouali, Imad, et al.
Published: (2022)
by: Aouali, Imad, et al.
Published: (2022)
RoiRL: Efficient, Self-Supervised Reasoning with Offline Iterative Reinforcement Learning
by: Arzhantsev, Aleksei, et al.
Published: (2025)
by: Arzhantsev, Aleksei, et al.
Published: (2025)
Prior-Dependent Allocations for Bayesian Fixed-Budget Best-Arm Identification in Structured Bandits
by: Nguyen, Nicolas, et al.
Published: (2024)
by: Nguyen, Nicolas, et al.
Published: (2024)
Sequential Off-Policy Learning with Logarithmic Smoothing
by: Haddouche, Maxime, et al.
Published: (2025)
by: Haddouche, Maxime, et al.
Published: (2025)
Unified PAC-Bayesian Study of Pessimism for Offline Policy Learning with Regularized Importance Sampling
by: Aouali, Imad, et al.
Published: (2024)
by: Aouali, Imad, et al.
Published: (2024)
Confounding is a Pervasive Problem in Real World Recommender Systems
by: Merkov, Alexander, et al.
Published: (2025)
by: Merkov, Alexander, et al.
Published: (2025)
A pragmatic policy learning approach to account for users' fatigue in repeated auctions
by: Heymann, Benjamin, et al.
Published: (2024)
by: Heymann, Benjamin, et al.
Published: (2024)
Fast Slate Policy Optimization: Going Beyond Plackett-Luce
by: Sakhi, Otmane, et al.
Published: (2023)
by: Sakhi, Otmane, et al.
Published: (2023)
Self-Consistency via Marginal Sharpening
by: Arzhantsev, Aleksei, et al.
Published: (2026)
by: Arzhantsev, Aleksei, et al.
Published: (2026)
Group-Sensitive Offline Contextual Bandits
by: Guo, Yihong, et al.
Published: (2025)
by: Guo, Yihong, et al.
Published: (2025)
Offline Contextual Bandits in the Presence of New Actions
by: Kishimoto, Ren, et al.
Published: (2026)
by: Kishimoto, Ren, et al.
Published: (2026)
Improved Offline Contextual Bandits with Second-Order Bounds: Betting and Freezing
by: Ryu, J. Jon, et al.
Published: (2025)
by: Ryu, J. Jon, et al.
Published: (2025)
Direction-Aware Offline-to-Online Learning in Linear Contextual Bandits
by: Han, Zean, et al.
Published: (2026)
by: Han, Zean, et al.
Published: (2026)
Leveraging Offline Data in Linear Latent Contextual Bandits
by: Kausik, Chinmaya, et al.
Published: (2024)
by: Kausik, Chinmaya, et al.
Published: (2024)
Bayesian Off-Policy Evaluation and Learning for Large Action Spaces
by: Aouali, Imad, et al.
Published: (2024)
by: Aouali, Imad, et al.
Published: (2024)
Thompson Sampling in Partially Observable Contextual Bandits
by: Park, Hongju, et al.
Published: (2024)
by: Park, Hongju, et al.
Published: (2024)
Thompson Sampling for Multi-Objective Linear Contextual Bandit
by: Park, Somangchan, et al.
Published: (2025)
by: Park, Somangchan, et al.
Published: (2025)
Upper Counterfactual Confidence Bounds: a New Optimism Principle for Contextual Bandits
by: Xu, Yunbei, et al.
Published: (2020)
by: Xu, Yunbei, et al.
Published: (2020)
The Sample Complexity of Multiclass and Sparse Contextual Bandits
by: Erez, Liad, et al.
Published: (2026)
by: Erez, Liad, et al.
Published: (2026)
Optimizing Warfarin Dosing Using Contextual Bandit: An Offline Policy Learning and Evaluation Method
by: Huang, Yong, et al.
Published: (2024)
by: Huang, Yong, et al.
Published: (2024)
Feel-Good Thompson Sampling for Contextual Dueling Bandits
by: Li, Xuheng, et al.
Published: (2024)
by: Li, Xuheng, et al.
Published: (2024)
Taming the Monster Every Context: Complexity Measure and Unified Framework for Offline-Oracle Efficient Contextual Bandits
by: Qin, Hao, et al.
Published: (2026)
by: Qin, Hao, et al.
Published: (2026)
Variance-Aware Feel-Good Thompson Sampling for Contextual Bandits
by: Li, Xuheng, et al.
Published: (2025)
by: Li, Xuheng, et al.
Published: (2025)
Provable Anytime Ensemble Sampling Algorithms in Nonlinear Contextual Bandits
by: Sun, Jiazheng, et al.
Published: (2025)
by: Sun, Jiazheng, et al.
Published: (2025)
The Sampling Complexity of Condorcet Winner Identification in Dueling Bandits
by: Saad, El Mehdi, et al.
Published: (2026)
by: Saad, El Mehdi, et al.
Published: (2026)
Fast and Sample Efficient Multi-Task Representation Learning in Stochastic Contextual Bandits
by: Lin, Jiabin, et al.
Published: (2024)
by: Lin, Jiabin, et al.
Published: (2024)
Bayesian Regret Minimization in Offline Bandits
by: Petrik, Marek, et al.
Published: (2023)
by: Petrik, Marek, et al.
Published: (2023)
Sparse Nonparametric Contextual Bandits
by: Flynn, Hamish, et al.
Published: (2025)
by: Flynn, Hamish, et al.
Published: (2025)
Budgeting Counterfactual for Offline RL
by: Liu, Yao, et al.
Published: (2023)
by: Liu, Yao, et al.
Published: (2023)
From Contextual Combinatorial Semi-Bandits to Bandit List Classification: Improved Sample Complexity with Sparse Rewards
by: Erez, Liad, et al.
Published: (2025)
by: Erez, Liad, et al.
Published: (2025)
Efficient Algorithms for Logistic Contextual Slate Bandits with Bandit Feedback
by: Goyal, Tanmay, et al.
Published: (2025)
by: Goyal, Tanmay, et al.
Published: (2025)
Offline Imitation Learning with Variational Counterfactual Reasoning
by: He, Bowei, et al.
Published: (2023)
by: He, Bowei, et al.
Published: (2023)
Similar Items
-
RecoAtlas: From Semantic Plausibility to Set-Level Utility in LLM Recommendation Agents
by: Aouali, Imad, et al.
Published: (2026) -
Off-Policy Learning in Large Action Spaces: Optimization Matters More Than Estimation
by: Aouali, Imad, et al.
Published: (2025) -
Non-Linear Counterfactual Aggregate Optimization
by: Heymann, Benjamin, et al.
Published: (2025) -
Exploiting Similarities in A/B Testing with Off-Policy Estimation
by: Sakhi, Otmane, et al.
Published: (2025) -
Learning to Bid in Repeated Second-Price Auctions with Dynamic Values and Aggregated Feedback
by: Heymann, Benjamin, et al.
Published: (2026)