Online Episodic Convex Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Moreno, Bianca Marin, Eldowa, Khaled, Gaillard, Pierre, Brégère, Margaux, Oudjane, Nadia |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MetaCURL: Non-stationary Concave Utility Reinforcement Learning
by: Moreno, Bianca Marin, et al.
Published: (2024)
by: Moreno, Bianca Marin, et al.
Published: (2024)
Online Markov Decision Processes with Terminal Law Constraints
by: Moreno, Bianca Marin, et al.
Published: (2026)
by: Moreno, Bianca Marin, et al.
Published: (2026)
Automated Spatio-Temporal Weather Modeling for Load Forecasting
by: Keisler, Julie, et al.
Published: (2024)
by: Keisler, Julie, et al.
Published: (2024)
AutoML Algorithms for Online Generalized Additive Model Selection: Application to Electricity Demand Forecasting
by: Das, Keshav, et al.
Published: (2025)
by: Das, Keshav, et al.
Published: (2025)
A Bandit Approach with Evolutionary Operators for Model Selection
by: Brégère, Margaux, et al.
Published: (2024)
by: Brégère, Margaux, et al.
Published: (2024)
Automated Deep Learning for Load Forecasting
by: Keisler, Julie, et al.
Published: (2024)
by: Keisler, Julie, et al.
Published: (2024)
Structured Prediction in Online Learning
by: Boudart, Pierre, et al.
Published: (2024)
by: Boudart, Pierre, et al.
Published: (2024)
Instance-Dependent Regret Bounds for Nonstochastic Linear Partial Monitoring
by: Di Gennaro, Federico, et al.
Published: (2025)
by: Di Gennaro, Federico, et al.
Published: (2025)
Sliding-Window Signatures for Time Series: Application to Electricity Demand Forecasting
by: Drobac, Nina, et al.
Published: (2025)
by: Drobac, Nina, et al.
Published: (2025)
Improved Regret Bounds for Bandits with Expert Advice
by: Cesa-Bianchi, Nicolò, et al.
Published: (2024)
by: Cesa-Bianchi, Nicolò, et al.
Published: (2024)
Information Capacity Regret Bounds for Bandits with Mediator Feedback
by: Eldowa, Khaled, et al.
Published: (2024)
by: Eldowa, Khaled, et al.
Published: (2024)
Minimax-optimal and Locally-adaptive Online Nonparametric Regression
by: Liautaud, Paul, et al.
Published: (2024)
by: Liautaud, Paul, et al.
Published: (2024)
Minimax Adaptive Online Nonparametric Regression over Besov Spaces
by: Liautaud, Paul, et al.
Published: (2025)
by: Liautaud, Paul, et al.
Published: (2025)
Online Learning Approach for Survival Analysis
by: Fernandez, Camila, et al.
Published: (2024)
by: Fernandez, Camila, et al.
Published: (2024)
High-Probability Minimax Adaptive Estimation in Besov Spaces via Online-to-Batch
by: Liautaud, Paul, et al.
Published: (2026)
by: Liautaud, Paul, et al.
Published: (2026)
Budgeted Online Active Learning with Expert Advice and Episodic Priors
by: Goebel, Kristen, et al.
Published: (2025)
by: Goebel, Kristen, et al.
Published: (2025)
Stop Relying on No-Choice and Do not Repeat the Moves: Optimal, Efficient and Practical Algorithms for Assortment Optimization
by: Saha, Aadirupa, et al.
Published: (2024)
by: Saha, Aadirupa, et al.
Published: (2024)
Reinforcement Learning from Multi-level and Episodic Human Feedback
by: Elahi, Muhammad Qasim, et al.
Published: (2025)
by: Elahi, Muhammad Qasim, et al.
Published: (2025)
From Generative to Episodic: Sample-Efficient Replicable Reinforcement Learning
by: Hopkins, Max, et al.
Published: (2025)
by: Hopkins, Max, et al.
Published: (2025)
Episodic Reinforcement Learning with Expanded State-reward Space
by: Liang, Dayang, et al.
Published: (2024)
by: Liang, Dayang, et al.
Published: (2024)
Diffusion-based Episodes Augmentation for Offline Multi-Agent Reinforcement Learning
by: Oh, Jihwan, et al.
Published: (2024)
by: Oh, Jihwan, et al.
Published: (2024)
MoRe-ERL: Learning Motion Residuals using Episodic Reinforcement Learning
by: Huang, Xi, et al.
Published: (2025)
by: Huang, Xi, et al.
Published: (2025)
Enjoying Non-linearity in Multinomial Logistic Bandits: A Minimax-Optimal Algorithm
by: Boudart, Pierre, et al.
Published: (2025)
by: Boudart, Pierre, et al.
Published: (2025)
Minimax Optimal Variance-Aware Regret Bounds for Multinomial Logistic MDPs
by: Boudart, Pierre, et al.
Published: (2026)
by: Boudart, Pierre, et al.
Published: (2026)
Combinatorial Multivariant Multi-Armed Bandits with Applications to Episodic Reinforcement Learning and Beyond
by: Liu, Xutong, et al.
Published: (2024)
by: Liu, Xutong, et al.
Published: (2024)
TOP-ERL: Transformer-based Off-Policy Episodic Reinforcement Learning
by: Li, Ge, et al.
Published: (2024)
by: Li, Ge, et al.
Published: (2024)
Convex Is Back: Solving Belief MDPs With Convexity-Informed Deep Reinforcement Learning
by: Koutas, Daniel, et al.
Published: (2025)
by: Koutas, Daniel, et al.
Published: (2025)
The Three Regimes of Offline-to-Online Reinforcement Learning
by: Li, Lu, et al.
Published: (2025)
by: Li, Lu, et al.
Published: (2025)
Efficient Episodic Memory Utilization of Cooperative Multi-Agent Reinforcement Learning
by: Na, Hyungho, et al.
Published: (2024)
by: Na, Hyungho, et al.
Published: (2024)
Latent Reward: LLM-Empowered Credit Assignment in Episodic Reinforcement Learning
by: Qu, Yun, et al.
Published: (2024)
by: Qu, Yun, et al.
Published: (2024)
Counterfactual Learning of Stochastic Policies with Continuous Actions
by: Zenati, Houssam, et al.
Published: (2020)
by: Zenati, Houssam, et al.
Published: (2020)
Posterior Sampling-based Online Learning for Episodic POMDPs
by: Tang, Dengwang, et al.
Published: (2023)
by: Tang, Dengwang, et al.
Published: (2023)
Episodic Future Thinking Mechanism for Multi-agent Reinforcement Learning
by: Lee, Dongsu, et al.
Published: (2024)
by: Lee, Dongsu, et al.
Published: (2024)
Quantum-Inspired Episode Selection for Monte Carlo Reinforcement Learning via QUBO Optimization
by: Salloum, Hadi, et al.
Published: (2026)
by: Salloum, Hadi, et al.
Published: (2026)
Grounding Large Language Models in Interactive Environments with Online Reinforcement Learning
by: Carta, Thomas, et al.
Published: (2023)
by: Carta, Thomas, et al.
Published: (2023)
Alternating Regret for Online Convex Optimization
by: Hait, Soumita, et al.
Published: (2025)
by: Hait, Soumita, et al.
Published: (2025)
Learning-Augmented Decentralized Online Convex Optimization in Networks
by: Li, Pengfei, et al.
Published: (2023)
by: Li, Pengfei, et al.
Published: (2023)
Online (Non-)Convex Learning via Tempered Optimism
by: Haddouche, Maxime, et al.
Published: (2023)
by: Haddouche, Maxime, et al.
Published: (2023)
Unsupervised-to-Online Reinforcement Learning
by: Kim, Junsu, et al.
Published: (2024)
by: Kim, Junsu, et al.
Published: (2024)
Open the Black Box: Step-based Policy Updates for Temporally-Correlated Episodic Reinforcement Learning
by: Li, Ge, et al.
Published: (2024)
by: Li, Ge, et al.
Published: (2024)
Similar Items
-
MetaCURL: Non-stationary Concave Utility Reinforcement Learning
by: Moreno, Bianca Marin, et al.
Published: (2024) -
Online Markov Decision Processes with Terminal Law Constraints
by: Moreno, Bianca Marin, et al.
Published: (2026) -
Automated Spatio-Temporal Weather Modeling for Load Forecasting
by: Keisler, Julie, et al.
Published: (2024) -
AutoML Algorithms for Online Generalized Additive Model Selection: Application to Electricity Demand Forecasting
by: Das, Keshav, et al.
Published: (2025) -
A Bandit Approach with Evolutionary Operators for Model Selection
by: Brégère, Margaux, et al.
Published: (2024)