Learning Upper Lower Value Envelopes to Shape Online RL: A Principled Approach
Fuente:
arXiv
Salvato in:
| Autori principali: | Reboul, Sebastian, Halconruy, Hélène, Douc, Randal |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Tessellation Localized Transfer learning for nonparametric regression
di: Halconruy, Hélène, et al.
Pubblicazione: (2026)
di: Halconruy, Hélène, et al.
Pubblicazione: (2026)
Laplace Transform Based Low-Complexity Learning of Continuous Markov Semigroups
di: Kostic, Vladimir R., et al.
Pubblicazione: (2024)
di: Kostic, Vladimir R., et al.
Pubblicazione: (2024)
When to Transfer: Adaptive Source Selection for Positive Transfer in Linear Models
di: Cherkaoui, Hamza, et al.
Pubblicazione: (2025)
di: Cherkaoui, Hamza, et al.
Pubblicazione: (2025)
On a Projection Least Squares Estimator for Jump Diffusion Processes
di: Halconruy, Hélène, et al.
Pubblicazione: (2022)
di: Halconruy, Hélène, et al.
Pubblicazione: (2022)
Optimal Lower Bounds for Online Multicalibration
di: Collina, Natalie, et al.
Pubblicazione: (2026)
di: Collina, Natalie, et al.
Pubblicazione: (2026)
Self-Organizing State-Space Models with Artificial Dynamics
di: Chen, Yuan, et al.
Pubblicazione: (2024)
di: Chen, Yuan, et al.
Pubblicazione: (2024)
On Stopping Times of Power-one Sequential Tests: Tight Lower and Upper Bounds
di: Agrawal, Shubhada, et al.
Pubblicazione: (2025)
di: Agrawal, Shubhada, et al.
Pubblicazione: (2025)
Upper Counterfactual Confidence Bounds: a New Optimism Principle for Contextual Bandits
di: Xu, Yunbei, et al.
Pubblicazione: (2020)
di: Xu, Yunbei, et al.
Pubblicazione: (2020)
An Upper Confidence Bound Approach to Estimating the Maximum Mean
di: Kun, Zhang, et al.
Pubblicazione: (2024)
di: Kun, Zhang, et al.
Pubblicazione: (2024)
Extrapolation in Statistical Learning with Extreme Value Theory
di: Engelke, Sebastian, et al.
Pubblicazione: (2026)
di: Engelke, Sebastian, et al.
Pubblicazione: (2026)
Set-Valued Policy Learning
di: Fuentes-Vicente, Laura, et al.
Pubblicazione: (2026)
di: Fuentes-Vicente, Laura, et al.
Pubblicazione: (2026)
Upper Bounds for Local Learning Coefficients of Three-Layer Neural Networks
di: Kurumadani, Yuki
Pubblicazione: (2026)
di: Kurumadani, Yuki
Pubblicazione: (2026)
Structured Prediction in Online Learning
di: Boudart, Pierre, et al.
Pubblicazione: (2024)
di: Boudart, Pierre, et al.
Pubblicazione: (2024)
Are First-Order Diffusion Samplers Really Slower? A Fast Forward-Value Approach
di: Jiao, Yuchen, et al.
Pubblicazione: (2025)
di: Jiao, Yuchen, et al.
Pubblicazione: (2025)
Distributionally-Constrained Adversaries in Online Learning
di: Blanchard, Moïse, et al.
Pubblicazione: (2025)
di: Blanchard, Moïse, et al.
Pubblicazione: (2025)
General Lower Bounds for Differentially Private Federated Learning with Arbitrary Public-Transcript Interactions
di: Li, Yicheng
Pubblicazione: (2026)
di: Li, Yicheng
Pubblicazione: (2026)
On the Natural Gradient of the Evidence Lower Bound
di: Ay, Nihat, et al.
Pubblicazione: (2023)
di: Ay, Nihat, et al.
Pubblicazione: (2023)
Learning Dynamic Bayesian Networks from Data: Foundations, First Principles and Numerical Comparisons
di: Kungurtsev, Vyacheslav, et al.
Pubblicazione: (2024)
di: Kungurtsev, Vyacheslav, et al.
Pubblicazione: (2024)
Risk Measures and Upper Probabilities: Coherence and Stratification
di: Fröhlich, Christian, et al.
Pubblicazione: (2022)
di: Fröhlich, Christian, et al.
Pubblicazione: (2022)
Lower Bounds on the Size of Markov Equivalence Classes
di: Jahn, Erik, et al.
Pubblicazione: (2025)
di: Jahn, Erik, et al.
Pubblicazione: (2025)
Online Selective Conformal Prediction with Asymmetric Rules: A Permutation Test Approach
di: Zheng, Mingyi, et al.
Pubblicazione: (2026)
di: Zheng, Mingyi, et al.
Pubblicazione: (2026)
A Gapped Scale-Sensitive Dimension and Lower Bounds for Offset Rademacher Complexity
di: Jia, Zeyu, et al.
Pubblicazione: (2025)
di: Jia, Zeyu, et al.
Pubblicazione: (2025)
Multimodal Bandits: Regret Lower Bounds and Optimal Algorithms
di: Réveillard, William, et al.
Pubblicazione: (2025)
di: Réveillard, William, et al.
Pubblicazione: (2025)
Principled Out-of-Distribution Generalization via Simplicity
di: Ge, Jiawei, et al.
Pubblicazione: (2025)
di: Ge, Jiawei, et al.
Pubblicazione: (2025)
Bayesian Design Principles for Frequentist Sequential Learning
di: Xu, Yunbei, et al.
Pubblicazione: (2023)
di: Xu, Yunbei, et al.
Pubblicazione: (2023)
Learning conditional distributions on continuous spaces
di: Bénézet, Cyril, et al.
Pubblicazione: (2024)
di: Bénézet, Cyril, et al.
Pubblicazione: (2024)
On the Generalization and Robustness in Conditional Value-at-Risk
di: Mulumudi, Dinesh Karthik, et al.
Pubblicazione: (2026)
di: Mulumudi, Dinesh Karthik, et al.
Pubblicazione: (2026)
MIST: Reliable Streaming Decision Trees for Online Class-Incremental Learning via McDiarmid Bound
di: Pham, Phu-Hoa, et al.
Pubblicazione: (2026)
di: Pham, Phu-Hoa, et al.
Pubblicazione: (2026)
Online Learning with Unknown Constraints
di: Sridharan, Karthik, et al.
Pubblicazione: (2024)
di: Sridharan, Karthik, et al.
Pubblicazione: (2024)
Information Geometry of Wasserstein Statistics on Shapes and Affine Deformations
di: Amari, Shun-ichi, et al.
Pubblicazione: (2023)
di: Amari, Shun-ichi, et al.
Pubblicazione: (2023)
Gaussian Process Upper Confidence Bounds in Distributed Point Target Tracking over Wireless Sensor Networks
di: Liu, Xingchi, et al.
Pubblicazione: (2024)
di: Liu, Xingchi, et al.
Pubblicazione: (2024)
The Condition-Number Principle for Prototype Clustering
di: Li, Romano, et al.
Pubblicazione: (2026)
di: Li, Romano, et al.
Pubblicazione: (2026)
Assouad, Fano, and Le Cam with Interaction: A Unifying Lower Bound Framework and Characterization for Bandit Learnability
di: Chen, Fan, et al.
Pubblicazione: (2024)
di: Chen, Fan, et al.
Pubblicazione: (2024)
Low-degree Lower bounds for clustering in moderate dimension
di: Carpentier, Alexandra, et al.
Pubblicazione: (2026)
di: Carpentier, Alexandra, et al.
Pubblicazione: (2026)
Doubly Robust Interval Estimation for Optimal Policy Evaluation in Online Learning
di: Shen, Ye, et al.
Pubblicazione: (2021)
di: Shen, Ye, et al.
Pubblicazione: (2021)
Online and Offline Robust Multivariate Linear Regression
di: Godichon-Baggioni, Antoine, et al.
Pubblicazione: (2024)
di: Godichon-Baggioni, Antoine, et al.
Pubblicazione: (2024)
Online Quantile Regression for Nonparametric Additive Models
di: Zhan, Haoran
Pubblicazione: (2026)
di: Zhan, Haoran
Pubblicazione: (2026)
The Good, the Bad, and the Sampled: a No-Regret Approach to Safe Online Classification
di: Baharav, Tavor Z., et al.
Pubblicazione: (2025)
di: Baharav, Tavor Z., et al.
Pubblicazione: (2025)
Minimax Optimality of Score-based Diffusion Models: Beyond the Density Lower Bound Assumptions
di: Zhang, Kaihong, et al.
Pubblicazione: (2024)
di: Zhang, Kaihong, et al.
Pubblicazione: (2024)
Online Performance Estimation with Unlabeled Data: A Bayesian Application of the Hui-Walter Paradigm
di: Slote, Kevin, et al.
Pubblicazione: (2024)
di: Slote, Kevin, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Tessellation Localized Transfer learning for nonparametric regression
di: Halconruy, Hélène, et al.
Pubblicazione: (2026) -
Laplace Transform Based Low-Complexity Learning of Continuous Markov Semigroups
di: Kostic, Vladimir R., et al.
Pubblicazione: (2024) -
When to Transfer: Adaptive Source Selection for Positive Transfer in Linear Models
di: Cherkaoui, Hamza, et al.
Pubblicazione: (2025) -
On a Projection Least Squares Estimator for Jump Diffusion Processes
di: Halconruy, Hélène, et al.
Pubblicazione: (2022) -
Optimal Lower Bounds for Online Multicalibration
di: Collina, Natalie, et al.
Pubblicazione: (2026)