Counterfactual Learning of Stochastic Policies with Continuous Actions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zenati, Houssam, Bietti, Alberto, Martin, Matthieu, Diemert, Eustache, Gaillard, Pierre, Mairal, Julien |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2020
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Doubly-Robust Estimation of Counterfactual Policy Mean Embeddings
von: Zenati, Houssam, et al.
Veröffentlicht: (2025)
von: Zenati, Houssam, et al.
Veröffentlicht: (2025)
Towards Efficient and Optimal Covariance-Adaptive Algorithms for Combinatorial Semi-Bandits
von: Zhou, Julien, et al.
Veröffentlicht: (2024)
von: Zhou, Julien, et al.
Veröffentlicht: (2024)
FairJob: A Real-World Dataset for Fairness in Online Systems
von: Vladimirova, Mariia, et al.
Veröffentlicht: (2024)
von: Vladimirova, Mariia, et al.
Veröffentlicht: (2024)
Double Debiased Machine Learning for Mediation Analysis with Continuous Treatments
von: Zenati, Houssam, et al.
Veröffentlicht: (2025)
von: Zenati, Houssam, et al.
Veröffentlicht: (2025)
Semiparametric Efficient Test for Interpretable Distributional Treatment Effects
von: Zenati, Houssam, et al.
Veröffentlicht: (2026)
von: Zenati, Houssam, et al.
Veröffentlicht: (2026)
Functional Natural Policy Gradients
von: Bibaut, Aurelien, et al.
Veröffentlicht: (2026)
von: Bibaut, Aurelien, et al.
Veröffentlicht: (2026)
Kernel Treatment Effects with Adaptively Collected Data
von: Zenati, Houssam, et al.
Veröffentlicht: (2025)
von: Zenati, Houssam, et al.
Veröffentlicht: (2025)
Logarithmic Regret for Unconstrained Submodular Maximization Stochastic Bandit
von: Zhou, Julien, et al.
Veröffentlicht: (2024)
von: Zhou, Julien, et al.
Veröffentlicht: (2024)
Density Ratio-Free Doubly Robust Proxy Causal Learning
von: Bozkurt, Bariscan, et al.
Veröffentlicht: (2025)
von: Bozkurt, Bariscan, et al.
Veröffentlicht: (2025)
Functional Bilevel Optimization for Machine Learning
von: Petrulionyte, Ieva, et al.
Veröffentlicht: (2024)
von: Petrulionyte, Ieva, et al.
Veröffentlicht: (2024)
Doubly Robust Proxy Causal Learning with Neural Mean Embeddings
von: Bozkurt, Bariscan, et al.
Veröffentlicht: (2026)
von: Bozkurt, Bariscan, et al.
Veröffentlicht: (2026)
Fast Best-in-Class Regret for Contextual Bandits
von: Girard, Samuel, et al.
Veröffentlicht: (2025)
von: Girard, Samuel, et al.
Veröffentlicht: (2025)
Counterfactual Explanations for Continuous Action Reinforcement Learning
von: Dong, Shuyang, et al.
Veröffentlicht: (2025)
von: Dong, Shuyang, et al.
Veröffentlicht: (2025)
Causal mediation analysis with one or multiple mediators: a comparative study
von: Abécassis, Judith, et al.
Veröffentlicht: (2025)
von: Abécassis, Judith, et al.
Veröffentlicht: (2025)
Semiparametric Efficient Bilevel Gradient Estimation
von: Khoury, Fares El, et al.
Veröffentlicht: (2026)
von: Khoury, Fares El, et al.
Veröffentlicht: (2026)
Structured Prediction in Online Learning
von: Boudart, Pierre, et al.
Veröffentlicht: (2024)
von: Boudart, Pierre, et al.
Veröffentlicht: (2024)
Learning Associative Memories with Gradient Descent
von: Cabannes, Vivien, et al.
Veröffentlicht: (2024)
von: Cabannes, Vivien, et al.
Veröffentlicht: (2024)
Assign and Add: A Mechanistic Study of Compositional Arithmetic
von: Exoo, Brady, et al.
Veröffentlicht: (2026)
von: Exoo, Brady, et al.
Veröffentlicht: (2026)
Efficient Inference after Directionally Stable Adaptive Experiments
von: Shen, Zikai, et al.
Veröffentlicht: (2026)
von: Shen, Zikai, et al.
Veröffentlicht: (2026)
Stop Relying on No-Choice and Do not Repeat the Moves: Optimal, Efficient and Practical Algorithms for Assortment Optimization
von: Saha, Aadirupa, et al.
Veröffentlicht: (2024)
von: Saha, Aadirupa, et al.
Veröffentlicht: (2024)
Learning to Recall with Transformers Beyond Orthogonal Embeddings
von: Vural, Nuri Mert, et al.
Veröffentlicht: (2026)
von: Vural, Nuri Mert, et al.
Veröffentlicht: (2026)
Learning Optimal Deterministic Policies with Stochastic Policy Gradients
von: Montenegro, Alessandro, et al.
Veröffentlicht: (2024)
von: Montenegro, Alessandro, et al.
Veröffentlicht: (2024)
Nonparametric Instrumental Variable Analysis Without Structural Equations: Debiased Inference on Functionals of Inverse Problems with No Solutions
von: Shen, Zikai, et al.
Veröffentlicht: (2026)
von: Shen, Zikai, et al.
Veröffentlicht: (2026)
Understanding the Mechanisms of Fast Hyperparameter Transfer
von: Ghosh, Nikhil, et al.
Veröffentlicht: (2025)
von: Ghosh, Nikhil, et al.
Veröffentlicht: (2025)
MAP Estimation with Denoisers: Convergence Rates and Guarantees
von: Pesme, Scott, et al.
Veröffentlicht: (2025)
von: Pesme, Scott, et al.
Veröffentlicht: (2025)
Challenges in Non-Polymeric Crystal Structure Prediction: Why a Geometric, Permutation-Invariant Loss is Needed
von: Jehanno, Emmanuel, et al.
Veröffentlicht: (2025)
von: Jehanno, Emmanuel, et al.
Veröffentlicht: (2025)
BAnG: Bidirectional Anchored Generation for Conditional RNA Design
von: Klypa, Roman, et al.
Veröffentlicht: (2025)
von: Klypa, Roman, et al.
Veröffentlicht: (2025)
Multimodal Alignment and Preference Optimization for Zero-Shot Conditional RNA Generation
von: Klypa, Roman, et al.
Veröffentlicht: (2026)
von: Klypa, Roman, et al.
Veröffentlicht: (2026)
Non-Stationary Functional Bilevel Optimization
von: Bohne, Jason, et al.
Veröffentlicht: (2026)
von: Bohne, Jason, et al.
Veröffentlicht: (2026)
Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory
von: Kim, Juno, et al.
Veröffentlicht: (2026)
von: Kim, Juno, et al.
Veröffentlicht: (2026)
Image Processing and Machine Learning for Hyperspectral Unmixing: An Overview and the HySUPP Python Package
von: Rasti, Behnood, et al.
Veröffentlicht: (2023)
von: Rasti, Behnood, et al.
Veröffentlicht: (2023)
Enjoying Non-linearity in Multinomial Logistic Bandits: A Minimax-Optimal Algorithm
von: Boudart, Pierre, et al.
Veröffentlicht: (2025)
von: Boudart, Pierre, et al.
Veröffentlicht: (2025)
Minimax Optimal Variance-Aware Regret Bounds for Multinomial Logistic MDPs
von: Boudart, Pierre, et al.
Veröffentlicht: (2026)
von: Boudart, Pierre, et al.
Veröffentlicht: (2026)
Level Set Teleportation: An Optimization Perspective
von: Mishkin, Aaron, et al.
Veröffentlicht: (2024)
von: Mishkin, Aaron, et al.
Veröffentlicht: (2024)
Actor-Accelerated Policy Dual Averaging for Reinforcement Learning in Continuous Action Spaces
von: Gao, Ji, et al.
Veröffentlicht: (2026)
von: Gao, Ji, et al.
Veröffentlicht: (2026)
Action-Adaptive Continual Learning: Enabling Policy Generalization under Dynamic Action Spaces
von: Pan, Chaofan, et al.
Veröffentlicht: (2025)
von: Pan, Chaofan, et al.
Veröffentlicht: (2025)
Distributional Associations vs In-Context Reasoning: A Study of Feed-forward and Attention Layers
von: Chen, Lei, et al.
Veröffentlicht: (2024)
von: Chen, Lei, et al.
Veröffentlicht: (2024)
Online Episodic Convex Reinforcement Learning
von: Moreno, Bianca Marin, et al.
Veröffentlicht: (2025)
von: Moreno, Bianca Marin, et al.
Veröffentlicht: (2025)
Minimax Adaptive Online Nonparametric Regression over Besov Spaces
von: Liautaud, Paul, et al.
Veröffentlicht: (2025)
von: Liautaud, Paul, et al.
Veröffentlicht: (2025)
Minimax-optimal and Locally-adaptive Online Nonparametric Regression
von: Liautaud, Paul, et al.
Veröffentlicht: (2024)
von: Liautaud, Paul, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Doubly-Robust Estimation of Counterfactual Policy Mean Embeddings
von: Zenati, Houssam, et al.
Veröffentlicht: (2025) -
Towards Efficient and Optimal Covariance-Adaptive Algorithms for Combinatorial Semi-Bandits
von: Zhou, Julien, et al.
Veröffentlicht: (2024) -
FairJob: A Real-World Dataset for Fairness in Online Systems
von: Vladimirova, Mariia, et al.
Veröffentlicht: (2024) -
Double Debiased Machine Learning for Mediation Analysis with Continuous Treatments
von: Zenati, Houssam, et al.
Veröffentlicht: (2025) -
Semiparametric Efficient Test for Interpretable Distributional Treatment Effects
von: Zenati, Houssam, et al.
Veröffentlicht: (2026)