Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Sakhi, Otmane, Aouali, Imad, Alquier, Pierre, Chopin, Nicolas |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Off-Policy Learning in Large Action Spaces: Optimization Matters More Than Estimation
by: Aouali, Imad, et al.
Published: (2025)
by: Aouali, Imad, et al.
Published: (2025)
Sequential Off-Policy Learning with Logarithmic Smoothing
by: Haddouche, Maxime, et al.
Published: (2025)
by: Haddouche, Maxime, et al.
Published: (2025)
Offline Contextual Bandit with Counterfactual Sample Identification
by: Gilotte, Alexandre, et al.
Published: (2025)
by: Gilotte, Alexandre, et al.
Published: (2025)
Fast Slate Policy Optimization: Going Beyond Plackett-Luce
by: Sakhi, Otmane, et al.
Published: (2023)
by: Sakhi, Otmane, et al.
Published: (2023)
Self-Consistency via Marginal Sharpening
by: Arzhantsev, Aleksei, et al.
Published: (2026)
by: Arzhantsev, Aleksei, et al.
Published: (2026)
RecoAtlas: From Semantic Plausibility to Set-Level Utility in LLM Recommendation Agents
by: Aouali, Imad, et al.
Published: (2026)
by: Aouali, Imad, et al.
Published: (2026)
Exploiting Similarities in A/B Testing with Off-Policy Estimation
by: Sakhi, Otmane, et al.
Published: (2025)
by: Sakhi, Otmane, et al.
Published: (2025)
Probabilistic Rank and Reward: A Scalable Model for Slate Recommendation
by: Aouali, Imad, et al.
Published: (2022)
by: Aouali, Imad, et al.
Published: (2022)
Bayesian Off-Policy Evaluation and Learning for Large Action Spaces
by: Aouali, Imad, et al.
Published: (2024)
by: Aouali, Imad, et al.
Published: (2024)
Learning to Bid in Repeated Second-Price Auctions with Dynamic Values and Aggregated Feedback
by: Heymann, Benjamin, et al.
Published: (2026)
by: Heymann, Benjamin, et al.
Published: (2026)
Diffusion Models Meet Contextual Bandits
by: Aouali, Imad
Published: (2024)
by: Aouali, Imad
Published: (2024)
Non-Linear Counterfactual Aggregate Optimization
by: Heymann, Benjamin, et al.
Published: (2025)
by: Heymann, Benjamin, et al.
Published: (2025)
RoiRL: Efficient, Self-Supervised Reasoning with Offline Iterative Reinforcement Learning
by: Arzhantsev, Aleksei, et al.
Published: (2025)
by: Arzhantsev, Aleksei, et al.
Published: (2025)
Pessimistic Off-Policy Optimization for Learning to Rank
by: Cief, Matej, et al.
Published: (2022)
by: Cief, Matej, et al.
Published: (2022)
Prior-Dependent Allocations for Bayesian Fixed-Budget Best-Arm Identification in Structured Bandits
by: Nguyen, Nicolas, et al.
Published: (2024)
by: Nguyen, Nicolas, et al.
Published: (2024)
Unified PAC-Bayesian Study of Pessimism for Offline Policy Learning with Regularized Importance Sampling
by: Aouali, Imad, et al.
Published: (2024)
by: Aouali, Imad, et al.
Published: (2024)
Data Valuation for LLM Fine-Tuning: Efficient Shapley Value Approximation via Language Model Arithmetic
by: Tamine, Mélissa, et al.
Published: (2025)
by: Tamine, Mélissa, et al.
Published: (2025)
User-friendly introduction to PAC-Bayes bounds
by: Alquier, Pierre
Published: (2021)
by: Alquier, Pierre
Published: (2021)
Minimax optimality of deep neural networks on dependent data via PAC-Bayes bounds
by: Alquier, Pierre, et al.
Published: (2024)
by: Alquier, Pierre, et al.
Published: (2024)
Empirical PAC-Bayes Bounds for Markov Chains
by: Karagulyan, Vahe, et al.
Published: (2025)
by: Karagulyan, Vahe, et al.
Published: (2025)
Variance-Aware Estimation of Kernel Mean Embedding
by: Wolfer, Geoffrey, et al.
Published: (2022)
by: Wolfer, Geoffrey, et al.
Published: (2022)
Pessimistic Risk-Aware Policy Learning in Contextual Bandits
by: Wan, Yilong, et al.
Published: (2026)
by: Wan, Yilong, et al.
Published: (2026)
Pessimistic Backward Policy for GFlowNets
by: Jang, Hyosoon, et al.
Published: (2024)
by: Jang, Hyosoon, et al.
Published: (2024)
Convergence of Statistical Estimators via Mutual Information Bounds
by: Khribch, El Mahdi, et al.
Published: (2024)
by: Khribch, El Mahdi, et al.
Published: (2024)
Robust Bayesian Inference via Variational Approximations of Generalized Rho-Posteriors
by: Khribch, EL Mahdi, et al.
Published: (2026)
by: Khribch, EL Mahdi, et al.
Published: (2026)
Bayes meets Bernstein at the Meta Level: an Analysis of Fast Rates in Meta-Learning with PAC-Bayes
by: Riou, Charles, et al.
Published: (2023)
by: Riou, Charles, et al.
Published: (2023)
Optimistic Model Rollouts for Pessimistic Offline Policy Optimization
by: Zhai, Yuanzhao, et al.
Published: (2024)
by: Zhai, Yuanzhao, et al.
Published: (2024)
Optimistic Policy Learning under Pessimistic Adversaries with Regret and Violation Guarantees
by: Ganguly, Sourav, et al.
Published: (2026)
by: Ganguly, Sourav, et al.
Published: (2026)
Finite sample properties of parametric MMD estimation: robustness to misspecification and dependence
by: Chérief-Abdellatif, Badr-Eddine, et al.
Published: (2019)
by: Chérief-Abdellatif, Badr-Eddine, et al.
Published: (2019)
Long-term Off-Policy Evaluation and Learning
by: Saito, Yuta, et al.
Published: (2024)
by: Saito, Yuta, et al.
Published: (2024)
Learning a Pessimistic Reward Model in RLHF
by: Xu, Yinglun, et al.
Published: (2025)
by: Xu, Yinglun, et al.
Published: (2025)
POLAR: A Pessimistic Model-based Policy Learning Algorithm for Dynamic Treatment Regimes
by: Zhang, Ruijia, et al.
Published: (2025)
by: Zhang, Ruijia, et al.
Published: (2025)
Off-Policy Evaluation and Learning for Matching Markets
by: Hayashi, Yudai, et al.
Published: (2025)
by: Hayashi, Yudai, et al.
Published: (2025)
Learning Action Embeddings for Off-Policy Evaluation
by: Cief, Matej, et al.
Published: (2023)
by: Cief, Matej, et al.
Published: (2023)
Log-Sum-Exponential Estimator for Off-Policy Evaluation and Learning
by: Behnamnia, Armin, et al.
Published: (2025)
by: Behnamnia, Armin, et al.
Published: (2025)
Cross-Validated Off-Policy Evaluation
by: Cief, Matej, et al.
Published: (2024)
by: Cief, Matej, et al.
Published: (2024)
Automated Off-Policy Estimator Selection via Supervised Learning
by: Felicioni, Nicolò, et al.
Published: (2024)
by: Felicioni, Nicolò, et al.
Published: (2024)
Off-Policy Evaluation for Ranking Policies under Deterministic Logging Policies
by: Tanaka, Koichi, et al.
Published: (2026)
by: Tanaka, Koichi, et al.
Published: (2026)
Context-Action Embedding Learning for Off-Policy Evaluation in Contextual Bandits
by: Chandak, Kushagra, et al.
Published: (2025)
by: Chandak, Kushagra, et al.
Published: (2025)
DOLCE: Decomposing Off-Policy Evaluation/Learning into Lagged and Current Effects
by: Tamano, Shu
Published: (2025)
by: Tamano, Shu
Published: (2025)
Similar Items
-
Off-Policy Learning in Large Action Spaces: Optimization Matters More Than Estimation
by: Aouali, Imad, et al.
Published: (2025) -
Sequential Off-Policy Learning with Logarithmic Smoothing
by: Haddouche, Maxime, et al.
Published: (2025) -
Offline Contextual Bandit with Counterfactual Sample Identification
by: Gilotte, Alexandre, et al.
Published: (2025) -
Fast Slate Policy Optimization: Going Beyond Plackett-Luce
by: Sakhi, Otmane, et al.
Published: (2023) -
Self-Consistency via Marginal Sharpening
by: Arzhantsev, Aleksei, et al.
Published: (2026)