Online Matching via Reinforcement Learning: An Expert Policy Orchestration Strategy
Fuente:
arXiv
Saved in:
| Main Authors: | Mignacco, Chiara, Jonckheere, Matthieu, Stoltz, Gilles |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Policy Optimization via Adv2: Adversarial Learning on Advantage Functions
by: Jonckheere, Matthieu, et al.
Published: (2023)
by: Jonckheere, Matthieu, et al.
Published: (2023)
Narrowing the Gap between Adversarial and Stochastic MDPs via Policy Optimization
by: Tiapkin, Daniil, et al.
Published: (2024)
by: Tiapkin, Daniil, et al.
Published: (2024)
Efficiency of Parallel and Restart Exploration Strategies in Model Free Stochastic Simulations
by: Garcia, Ernesto, et al.
Published: (2025)
by: Garcia, Ernesto, et al.
Published: (2025)
Evolving Diffusion and Flow Matching Policies for Online Reinforcement Learning
by: Zhang, Chubin, et al.
Published: (2025)
by: Zhang, Chubin, et al.
Published: (2025)
Optimization Trade-offs in Asynchronous Federated Learning: A Stochastic Networks Approach
by: Alahyane, Abdelkrim, et al.
Published: (2026)
by: Alahyane, Abdelkrim, et al.
Published: (2026)
Score-Aware Policy-Gradient and Performance Guarantees using Local Lyapunov Stability
by: Comte, Céline, et al.
Published: (2023)
by: Comte, Céline, et al.
Published: (2023)
Blackwell's Approachability for Sequential Conformal Inference
by: Principato, Guillaume, et al.
Published: (2025)
by: Principato, Guillaume, et al.
Published: (2025)
Queuing dynamics of asynchronous Federated Learning
by: Leconte, Louis, et al.
Published: (2024)
by: Leconte, Louis, et al.
Published: (2024)
ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning
by: Zhang, Tonghe, et al.
Published: (2025)
by: Zhang, Tonghe, et al.
Published: (2025)
Reinforcement Learning for Flow-Matching Policies
by: Pfrommer, Samuel, et al.
Published: (2025)
by: Pfrommer, Samuel, et al.
Published: (2025)
Optimizing Asynchronous Federated Learning: A Delicate Trade-Off Between Model-Parameter Staleness and Update Frequency
by: Alahyane, Abdelkrim, et al.
Published: (2025)
by: Alahyane, Abdelkrim, et al.
Published: (2025)
Parametrized Power-Iteration Clustering for Directed Graphs
by: Debaussart-Joniec, Gwendal, et al.
Published: (2022)
by: Debaussart-Joniec, Gwendal, et al.
Published: (2022)
Optimal Protocols for Continual Learning via Statistical Physics and Control Theory
by: Mori, Francesco, et al.
Published: (2024)
by: Mori, Francesco, et al.
Published: (2024)
Efficient Online Reinforcement Learning for Diffusion Policy
by: Ma, Haitong, et al.
Published: (2025)
by: Ma, Haitong, et al.
Published: (2025)
Controllable Flow Matching for Online Reinforcement Learning
by: Wang, Bin, et al.
Published: (2025)
by: Wang, Bin, et al.
Published: (2025)
Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies
by: Li, Zeyang, et al.
Published: (2026)
by: Li, Zeyang, et al.
Published: (2026)
Pareto Inverse Reinforcement Learning for Diverse Expert Policy Generation
by: Kim, Woo Kyung, et al.
Published: (2024)
by: Kim, Woo Kyung, et al.
Published: (2024)
Diversity-Preserving K-Armed Bandits, Revisited
by: Hadiji, Hédi, et al.
Published: (2020)
by: Hadiji, Hédi, et al.
Published: (2020)
Analytic theory of dropout regularization
by: Mori, Francesco, et al.
Published: (2025)
by: Mori, Francesco, et al.
Published: (2025)
Discrete Flow Matching for Offline-to-Online Reinforcement Learning
by: Khan, Fairoz Nower, et al.
Published: (2026)
by: Khan, Fairoz Nower, et al.
Published: (2026)
Generalized Dirichlet Energy and Graph Laplacians for Clustering Directed and Undirected Graphs
by: Sevi, Harry, et al.
Published: (2022)
by: Sevi, Harry, et al.
Published: (2022)
Towards Fast Safe Online Reinforcement Learning via Policy Finetuning
by: Chen, Keru, et al.
Published: (2024)
by: Chen, Keru, et al.
Published: (2024)
Deep Reinforcement Learning for Online Optimal Execution Strategies
by: Micheli, Alessandro, et al.
Published: (2024)
by: Micheli, Alessandro, et al.
Published: (2024)
Flow-Based Policy for Online Reinforcement Learning
by: Lv, Lei, et al.
Published: (2025)
by: Lv, Lei, et al.
Published: (2025)
Active Reinforcement Learning Strategies for Offline Policy Improvement
by: Dukkipati, Ambedkar, et al.
Published: (2024)
by: Dukkipati, Ambedkar, et al.
Published: (2024)
Online Estimation and Inference for Robust Policy Evaluation in Reinforcement Learning
by: Liu, Weidong, et al.
Published: (2023)
by: Liu, Weidong, et al.
Published: (2023)
Prism: Policy Reuse via Interpretable Strategy Mapping in Reinforcement Learning
by: Pravetz, Thomas
Published: (2026)
by: Pravetz, Thomas
Published: (2026)
Flow Matching with Injected Noise for Offline-to-Online Reinforcement Learning
by: Shin, Yongjae, et al.
Published: (2026)
by: Shin, Yongjae, et al.
Published: (2026)
A statistical physics framework for optimal learning
by: Mignacco, Francesca, et al.
Published: (2025)
by: Mignacco, Francesca, et al.
Published: (2025)
RLOMM: An Efficient and Robust Online Map Matching Framework with Reinforcement Learning
by: Chen, Minxiao, et al.
Published: (2025)
by: Chen, Minxiao, et al.
Published: (2025)
Minimax Optimal Strategy for Delayed Observations in Online Reinforcement Learning
by: Lee, Harin, et al.
Published: (2026)
by: Lee, Harin, et al.
Published: (2026)
On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting
by: Zhang, Wenhao, et al.
Published: (2025)
by: Zhang, Wenhao, et al.
Published: (2025)
FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning
by: Wan, Zhenglin, et al.
Published: (2025)
by: Wan, Zhenglin, et al.
Published: (2025)
Iterative Refinement of Flow Policies in Probability Space for Online Reinforcement Learning
by: Sun, Mingyang, et al.
Published: (2025)
by: Sun, Mingyang, et al.
Published: (2025)
A Non-Monolithic Policy Approach of Offline-to-Online Reinforcement Learning
by: Kim, JaeYoon, et al.
Published: (2024)
by: Kim, JaeYoon, et al.
Published: (2024)
Improving Policy Exploitation in Online Reinforcement Learning with Instant Retrospect Action
by: Gao, Gong, et al.
Published: (2026)
by: Gao, Gong, et al.
Published: (2026)
Expert-Free Online Transfer Learning in Multi-Agent Reinforcement Learning
by: Castagna, Alberto
Published: (2025)
by: Castagna, Alberto
Published: (2025)
Online Learning-to-Defer with Varying Experts
by: Duy, Dang Hoang, et al.
Published: (2026)
by: Duy, Dang Hoang, et al.
Published: (2026)
Learning by Doing: An Online Causal Reinforcement Learning Framework with Causal-Aware Policy
by: Cai, Ruichu, et al.
Published: (2024)
by: Cai, Ruichu, et al.
Published: (2024)
Score and Distribution Matching Policy: Advanced Accelerated Visuomotor Policies via Matched Distillation
by: Jia, Bofang, et al.
Published: (2024)
by: Jia, Bofang, et al.
Published: (2024)
Similar Items
-
Policy Optimization via Adv2: Adversarial Learning on Advantage Functions
by: Jonckheere, Matthieu, et al.
Published: (2023) -
Narrowing the Gap between Adversarial and Stochastic MDPs via Policy Optimization
by: Tiapkin, Daniil, et al.
Published: (2024) -
Efficiency of Parallel and Restart Exploration Strategies in Model Free Stochastic Simulations
by: Garcia, Ernesto, et al.
Published: (2025) -
Evolving Diffusion and Flow Matching Policies for Online Reinforcement Learning
by: Zhang, Chubin, et al.
Published: (2025) -
Optimization Trade-offs in Asynchronous Federated Learning: A Stochastic Networks Approach
by: Alahyane, Abdelkrim, et al.
Published: (2026)