K-Myriad: Jump-starting reinforcement learning with unsupervised parallel agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | De Paola, Vincenzo, Mutti, Mirco, Zamboni, Riccardo, Restelli, Marcello |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Enhancing Diversity in Parallel Agents: A Maximum State Entropy Exploration Story
von: De Paola, Vincenzo, et al.
Veröffentlicht: (2025)
von: De Paola, Vincenzo, et al.
Veröffentlicht: (2025)
Towards Principled Unsupervised Multi-Agent Reinforcement Learning
von: Zamboni, Riccardo, et al.
Veröffentlicht: (2025)
von: Zamboni, Riccardo, et al.
Veröffentlicht: (2025)
The Limits of Pure Exploration in POMDPs: When the Observation Entropy is Enough
von: Zamboni, Riccardo, et al.
Veröffentlicht: (2024)
von: Zamboni, Riccardo, et al.
Veröffentlicht: (2024)
How to Explore with Belief: State Entropy Maximization in POMDPs
von: Zamboni, Riccardo, et al.
Veröffentlicht: (2024)
von: Zamboni, Riccardo, et al.
Veröffentlicht: (2024)
From Parameters to Behaviors: Unsupervised Compression of the Policy Space
von: Tenedini, Davide, et al.
Veröffentlicht: (2025)
von: Tenedini, Davide, et al.
Veröffentlicht: (2025)
Unsupervised Behavioral Compression: Learning Low-Dimensional Policy Manifolds through State-Occupancy Matching
von: Fraschini, Andrea, et al.
Veröffentlicht: (2026)
von: Fraschini, Andrea, et al.
Veröffentlicht: (2026)
Exploiting Causal Graph Priors with Posterior Sampling for Reinforcement Learning
von: Mutti, Mirco, et al.
Veröffentlicht: (2023)
von: Mutti, Mirco, et al.
Veröffentlicht: (2023)
Scalable Multi-Agent Offline Reinforcement Learning and the Role of Information
von: Zamboni, Riccardo, et al.
Veröffentlicht: (2025)
von: Zamboni, Riccardo, et al.
Veröffentlicht: (2025)
Test-Time Regret Minimization in Meta Reinforcement Learning
von: Mutti, Mirco, et al.
Veröffentlicht: (2024)
von: Mutti, Mirco, et al.
Veröffentlicht: (2024)
"So, Tell Me About Your Policy...": Distillation of interpretable policies from Deep Reinforcement Learning agents
von: Dispoto, Giovanni, et al.
Veröffentlicht: (2025)
von: Dispoto, Giovanni, et al.
Veröffentlicht: (2025)
Pure Exploration under Mediators' Feedback
von: Poiani, Riccardo, et al.
Veröffentlicht: (2023)
von: Poiani, Riccardo, et al.
Veröffentlicht: (2023)
Building surrogate models using trajectories of agents trained by Reinforcement Learning
von: Cestero, Julen, et al.
Veröffentlicht: (2025)
von: Cestero, Julen, et al.
Veröffentlicht: (2025)
Online Market Making and the Value of Observing the Order Book
von: Maran, Davide, et al.
Veröffentlicht: (2026)
von: Maran, Davide, et al.
Veröffentlicht: (2026)
Online Dynamic Pricing of Complementary Products
von: Mussi, Marco, et al.
Veröffentlicht: (2025)
von: Mussi, Marco, et al.
Veröffentlicht: (2025)
Finite Sample Bounds for Non-Parametric Regression: Optimal Sample Efficiency and Space Complexity
von: Maran, Davide, et al.
Veröffentlicht: (2024)
von: Maran, Davide, et al.
Veröffentlicht: (2024)
Geometric Active Exploration in Markov Decision Processes: the Benefit of Abstraction
von: De Santi, Riccardo, et al.
Veröffentlicht: (2024)
von: De Santi, Riccardo, et al.
Veröffentlicht: (2024)
Reward Compatibility: A Framework for Inverse RL
von: Lazzati, Filippo, et al.
Veröffentlicht: (2025)
von: Lazzati, Filippo, et al.
Veröffentlicht: (2025)
Truncating Trajectories in Monte Carlo Policy Evaluation: an Adaptive Approach
von: Poiani, Riccardo, et al.
Veröffentlicht: (2024)
von: Poiani, Riccardo, et al.
Veröffentlicht: (2024)
Offline Inverse RL: New Solution Concepts and Provably Efficient Algorithms
von: Lazzati, Filippo, et al.
Veröffentlicht: (2024)
von: Lazzati, Filippo, et al.
Veröffentlicht: (2024)
How does Inverse RL Scale to Large State Spaces? A Provably Efficient Approach
von: Lazzati, Filippo, et al.
Veröffentlicht: (2024)
von: Lazzati, Filippo, et al.
Veröffentlicht: (2024)
Inverse Reinforcement Learning with Sub-optimal Experts
von: Poiani, Riccardo, et al.
Veröffentlicht: (2024)
von: Poiani, Riccardo, et al.
Veröffentlicht: (2024)
Learning in Markov Decision Processes with Exogenous Dynamics
von: Maran, Davide, et al.
Veröffentlicht: (2026)
von: Maran, Davide, et al.
Veröffentlicht: (2026)
Optimal Multi-Fidelity Best-Arm Identification
von: Poiani, Riccardo, et al.
Veröffentlicht: (2024)
von: Poiani, Riccardo, et al.
Veröffentlicht: (2024)
A Classification View on Meta Learning Bandits
von: Mutti, Mirco, et al.
Veröffentlicht: (2025)
von: Mutti, Mirco, et al.
Veröffentlicht: (2025)
Achieving $\widetilde{\mathcal{O}}(\sqrt{T})$ Regret in Average-Reward POMDPs with Known Observation Models
von: Russo, Alessio, et al.
Veröffentlicht: (2025)
von: Russo, Alessio, et al.
Veröffentlicht: (2025)
Interpetable Target-Feature Aggregation for Multi-Task Learning based on Bias-Variance Analysis
von: Bonetti, Paolo, et al.
Veröffentlicht: (2024)
von: Bonetti, Paolo, et al.
Veröffentlicht: (2024)
Efficient Learning of POMDPs with Known Observation Model in Average-Reward Setting
von: Russo, Alessio, et al.
Veröffentlicht: (2024)
von: Russo, Alessio, et al.
Veröffentlicht: (2024)
A Provably Efficient Option-Based Algorithm for both High-Level and Low-Level Learning
von: Drappo, Gianluca, et al.
Veröffentlicht: (2024)
von: Drappo, Gianluca, et al.
Veröffentlicht: (2024)
How Log-Barrier Helps Exploration in Policy Optimization
von: Cesani, Leonardo, et al.
Veröffentlicht: (2026)
von: Cesani, Leonardo, et al.
Veröffentlicht: (2026)
A Theoretical Framework for Partially Observed Reward-States in RLHF
von: Kausik, Chinmaya, et al.
Veröffentlicht: (2024)
von: Kausik, Chinmaya, et al.
Veröffentlicht: (2024)
Temporal Difference Calibration in Sequential Tasks: Application to Vision-Language-Action Models
von: Francis-Meretzki, Shelly, et al.
Veröffentlicht: (2026)
von: Francis-Meretzki, Shelly, et al.
Veröffentlicht: (2026)
Large-scale automatic carbon ion treatment planning for head and neck cancers via parallel multi-agent reinforcement learning
von: Zhang, Jueye, et al.
Veröffentlicht: (2025)
von: Zhang, Jueye, et al.
Veröffentlicht: (2025)
Local Linearity: the Key for No-regret Reinforcement Learning in Continuous MDPs
von: Maran, Davide, et al.
Veröffentlicht: (2024)
von: Maran, Davide, et al.
Veröffentlicht: (2024)
Policy Gradient with Active Importance Sampling
von: Papini, Matteo, et al.
Veröffentlicht: (2024)
von: Papini, Matteo, et al.
Veröffentlicht: (2024)
A Reinforcement Learning Approach for Optimal Control in Microgrids
von: Salaorni, Davide, et al.
Veröffentlicht: (2025)
von: Salaorni, Davide, et al.
Veröffentlicht: (2025)
ModelLens: Finding the Best for Your Task from Myriads of Models
von: Cai, Rui, et al.
Veröffentlicht: (2026)
von: Cai, Rui, et al.
Veröffentlicht: (2026)
Blindfolded Experts Generalize Better: Insights from Robotic Manipulation and Videogames
von: Zisselman, Ev, et al.
Veröffentlicht: (2025)
von: Zisselman, Ev, et al.
Veröffentlicht: (2025)
parallelcbf: A composable safety-filter and auditability framework for tensor-parallel reinforcement learning
von: Lu, Yijun, et al.
Veröffentlicht: (2026)
von: Lu, Yijun, et al.
Veröffentlicht: (2026)
Information Capacity Regret Bounds for Bandits with Mediator Feedback
von: Eldowa, Khaled, et al.
Veröffentlicht: (2024)
von: Eldowa, Khaled, et al.
Veröffentlicht: (2024)
AI Pangaea: Unifying Intelligence Islands for Adapting Myriad Tasks
von: Chang, Jianlong, et al.
Veröffentlicht: (2025)
von: Chang, Jianlong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Enhancing Diversity in Parallel Agents: A Maximum State Entropy Exploration Story
von: De Paola, Vincenzo, et al.
Veröffentlicht: (2025) -
Towards Principled Unsupervised Multi-Agent Reinforcement Learning
von: Zamboni, Riccardo, et al.
Veröffentlicht: (2025) -
The Limits of Pure Exploration in POMDPs: When the Observation Entropy is Enough
von: Zamboni, Riccardo, et al.
Veröffentlicht: (2024) -
How to Explore with Belief: State Entropy Maximization in POMDPs
von: Zamboni, Riccardo, et al.
Veröffentlicht: (2024) -
From Parameters to Behaviors: Unsupervised Compression of the Policy Space
von: Tenedini, Davide, et al.
Veröffentlicht: (2025)