Sample Complexity Bounds for Stochastic Shortest Path with a Generative Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tarbouriech, Jean, Pirotta, Matteo, Valko, Michal, Lazaric, Alessandro |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Large-scale semi-supervised learning with online spectral graph sparsification
von: Calandriello, Daniele, et al.
Veröffentlicht: (2026)
von: Calandriello, Daniele, et al.
Veröffentlicht: (2026)
Analysis of Nystrom method with sequential ridge leverage scores
von: Calandriello, Daniele, et al.
Veröffentlicht: (2026)
von: Calandriello, Daniele, et al.
Veröffentlicht: (2026)
Pack only the essentials: Adaptive dictionary learning for kernel ridge regression
von: Calandriello, Daniele, et al.
Veröffentlicht: (2026)
von: Calandriello, Daniele, et al.
Veröffentlicht: (2026)
Maximum Entropy Semi-Supervised Inverse Reinforcement Learning
von: Audiffren, Julien, et al.
Veröffentlicht: (2026)
von: Audiffren, Julien, et al.
Veröffentlicht: (2026)
A single algorithm for both restless and rested rotting bandits
von: Seznec, Julien, et al.
Veröffentlicht: (2026)
von: Seznec, Julien, et al.
Veröffentlicht: (2026)
Improved large-scale graph learning through ridge spectral sparsification
von: Calandriello, Daniele, et al.
Veröffentlicht: (2026)
von: Calandriello, Daniele, et al.
Veröffentlicht: (2026)
TD-JEPA: Latent-predictive Representations for Zero-Shot Reinforcement Learning
von: Bagatella, Marco, et al.
Veröffentlicht: (2025)
von: Bagatella, Marco, et al.
Veröffentlicht: (2025)
Trading off rewards and errors in multi-armed bandits
von: Erraqabi, Akram, et al.
Veröffentlicht: (2026)
von: Erraqabi, Akram, et al.
Veröffentlicht: (2026)
Compositional Planning with Jumpy World Models
von: Farebrother, Jesse, et al.
Veröffentlicht: (2026)
von: Farebrother, Jesse, et al.
Veröffentlicht: (2026)
Simple Ingredients for Offline Reinforcement Learning
von: Cetin, Edoardo, et al.
Veröffentlicht: (2024)
von: Cetin, Edoardo, et al.
Veröffentlicht: (2024)
Temporal Difference Flows
von: Farebrother, Jesse, et al.
Veröffentlicht: (2025)
von: Farebrother, Jesse, et al.
Veröffentlicht: (2025)
Zero-Shot Whole-Body Humanoid Control via Behavioral Foundation Models
von: Tirinzoni, Andrea, et al.
Veröffentlicht: (2025)
von: Tirinzoni, Andrea, et al.
Veröffentlicht: (2025)
Blazing the trails before beating the path: Sample-efficient Monte-Carlo planning
von: Grill, Jean-Bastien, et al.
Veröffentlicht: (2026)
von: Grill, Jean-Bastien, et al.
Veröffentlicht: (2026)
Stochastic simultaneous optimistic optimization
von: Valko, Michal, et al.
Veröffentlicht: (2026)
von: Valko, Michal, et al.
Veröffentlicht: (2026)
Regret Lower Bounds for Decentralized Multi-Agent Stochastic Shortest Path Problems
von: Chavan, Utkarsh U., et al.
Veröffentlicht: (2025)
von: Chavan, Utkarsh U., et al.
Veröffentlicht: (2025)
Bandits on graphs and structures
von: Valko, Michal
Veröffentlicht: (2026)
von: Valko, Michal
Veröffentlicht: (2026)
Adaptive graph-based algorithms for conditional anomaly detection and semi-supervised learning
von: Valko, Michal
Veröffentlicht: (2026)
von: Valko, Michal
Veröffentlicht: (2026)
Fast Adaptation with Behavioral Foundation Models
von: Sikchi, Harshit, et al.
Veröffentlicht: (2025)
von: Sikchi, Harshit, et al.
Veröffentlicht: (2025)
Stochastic Shortest Path with Sparse Adversarial Costs
von: Johnson, Emmeran, et al.
Veröffentlicht: (2025)
von: Johnson, Emmeran, et al.
Veröffentlicht: (2025)
Regret Guarantees for Linear Contextual Stochastic Shortest Path
von: Polikar, Dor, et al.
Veröffentlicht: (2025)
von: Polikar, Dor, et al.
Veröffentlicht: (2025)
Language Generation with Replay: A Learning-Theoretic View of Model Collapse
von: Racca, Giorgio, et al.
Veröffentlicht: (2026)
von: Racca, Giorgio, et al.
Veröffentlicht: (2026)
Convergent Reinforcement Learning Algorithms for Stochastic Shortest Path Problem
von: Guin, Soumyajit, et al.
Veröffentlicht: (2025)
von: Guin, Soumyajit, et al.
Veröffentlicht: (2025)
Learning from a single labeled face and a stream of unlabeled data
von: Kveton, Branislav, et al.
Veröffentlicht: (2026)
von: Kveton, Branislav, et al.
Veröffentlicht: (2026)
Feature importance analysis for patient management decisions
von: Valko, Michal, et al.
Veröffentlicht: (2026)
von: Valko, Michal, et al.
Veröffentlicht: (2026)
Online combinatorial optimization with stochastic decision sets and adversarial losses
von: Neu, Gergely, et al.
Veröffentlicht: (2026)
von: Neu, Gergely, et al.
Veröffentlicht: (2026)
Distance metric learning for conditional anomaly detection
von: Valko, Michal, et al.
Veröffentlicht: (2026)
von: Valko, Michal, et al.
Veröffentlicht: (2026)
Revealing graph bandits for maximizing local influence
von: Carpentier, Alexandra, et al.
Veröffentlicht: (2026)
von: Carpentier, Alexandra, et al.
Veröffentlicht: (2026)
Extreme bandits
von: Carpentier, Alexandra, et al.
Veröffentlicht: (2026)
von: Carpentier, Alexandra, et al.
Veröffentlicht: (2026)
Black-box optimization of noisy functions with unknown smoothness
von: Grill, Jean-Bastien, et al.
Veröffentlicht: (2026)
von: Grill, Jean-Bastien, et al.
Veröffentlicht: (2026)
Reinforcement Learning with Options and State Representation
von: Ghriss, Ayoub, et al.
Veröffentlicht: (2024)
von: Ghriss, Ayoub, et al.
Veröffentlicht: (2024)
Nearly Minimax Optimal Regret for Learning Linear Mixture Stochastic Shortest Path
von: Di, Qiwei, et al.
Veröffentlicht: (2024)
von: Di, Qiwei, et al.
Veröffentlicht: (2024)
The Harder Path: Last Iterate Convergence for Uncoupled Learning in Zero-Sum Games with Bandit Feedback
von: Fiegel, Côme, et al.
Veröffentlicht: (2026)
von: Fiegel, Côme, et al.
Veröffentlicht: (2026)
Learning Shortest Paths with Generative Flow Networks
von: Morozov, Nikita, et al.
Veröffentlicht: (2026)
von: Morozov, Nikita, et al.
Veröffentlicht: (2026)
Efficient Estimation of Shortest-Path Distance Distributions to Samples in Graphs
von: Zhu, Alan, et al.
Veröffentlicht: (2025)
von: Zhu, Alan, et al.
Veröffentlicht: (2025)
Sample Complexity Bounds for Linear Constrained MDPs with a Generative Model
von: Liu, Xingtu, et al.
Veröffentlicht: (2025)
von: Liu, Xingtu, et al.
Veröffentlicht: (2025)
Generalization in LLM Problem Solving: The Case of the Shortest Path
von: Tong, Yao, et al.
Veröffentlicht: (2026)
von: Tong, Yao, et al.
Veröffentlicht: (2026)
Best of both worlds: Stochastic & adversarial best-arm identification
von: Abbasi-Yadkori, Yasin, et al.
Veröffentlicht: (2026)
von: Abbasi-Yadkori, Yasin, et al.
Veröffentlicht: (2026)
Active multiple matrix completion with adaptive confidence sets
von: Locatelli, Andrea, et al.
Veröffentlicht: (2026)
von: Locatelli, Andrea, et al.
Veröffentlicht: (2026)
Bandits attack function optimization
von: Preux, Philippe, et al.
Veröffentlicht: (2026)
von: Preux, Philippe, et al.
Veröffentlicht: (2026)
Online learning with Erdős-Rényi side-observation graphs
von: Kocák, Tomáš, et al.
Veröffentlicht: (2026)
von: Kocák, Tomáš, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Large-scale semi-supervised learning with online spectral graph sparsification
von: Calandriello, Daniele, et al.
Veröffentlicht: (2026) -
Analysis of Nystrom method with sequential ridge leverage scores
von: Calandriello, Daniele, et al.
Veröffentlicht: (2026) -
Pack only the essentials: Adaptive dictionary learning for kernel ridge regression
von: Calandriello, Daniele, et al.
Veröffentlicht: (2026) -
Maximum Entropy Semi-Supervised Inverse Reinforcement Learning
von: Audiffren, Julien, et al.
Veröffentlicht: (2026) -
A single algorithm for both restless and rested rotting bandits
von: Seznec, Julien, et al.
Veröffentlicht: (2026)