Trajectory Data Suffices for Statistically Efficient Learning in Offline RL with Linear $q^π$-Realizability and Concentrability
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tkachuk, Volodymyr, Weisz, Gellért, Szepesvári, Csaba |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Regret Minimization via Saddle Point Optimization
von: Kirschner, Johannes, et al.
Veröffentlicht: (2024)
von: Kirschner, Johannes, et al.
Veröffentlicht: (2024)
Confident Natural Policy Gradient for Local Planning in $q_π$-realizable Constrained MDPs
von: Tian, Tian, et al.
Veröffentlicht: (2024)
von: Tian, Tian, et al.
Veröffentlicht: (2024)
Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence
von: György, András, et al.
Veröffentlicht: (2025)
von: György, András, et al.
Veröffentlicht: (2025)
Rectifying Regression in Reinforcement Learning
von: Ayoub, Alex, et al.
Veröffentlicht: (2025)
von: Ayoub, Alex, et al.
Veröffentlicht: (2025)
Sharper Guarantees for Misspecified Kernelized Bandit Optimization
von: Maran, Davide, et al.
Veröffentlicht: (2026)
von: Maran, Davide, et al.
Veröffentlicht: (2026)
Learning to Reason Efficiently with Discounted Reinforcement Learning
von: Ayoub, Alex, et al.
Veröffentlicht: (2025)
von: Ayoub, Alex, et al.
Veröffentlicht: (2025)
Frozen Policy Iteration: Computationally Efficient RL under Linear $Q^π$ Realizability for Deterministic Dynamics
von: Ke, Yijing, et al.
Veröffentlicht: (2026)
von: Ke, Yijing, et al.
Veröffentlicht: (2026)
Switching the Loss Reduces the Cost in Batch (Offline) Reinforcement Learning
von: Ayoub, Alex, et al.
Veröffentlicht: (2024)
von: Ayoub, Alex, et al.
Veröffentlicht: (2024)
Sharp analysis of linear ensemble sampling
von: Akhavan, Arya, et al.
Veröffentlicht: (2026)
von: Akhavan, Arya, et al.
Veröffentlicht: (2026)
Investigating Action Encodings in Recurrent Neural Networks in Reinforcement Learning
von: Schlegel, Matthew, et al.
Veröffentlicht: (2026)
von: Schlegel, Matthew, et al.
Veröffentlicht: (2026)
Balancing optimism and pessimism in offline-to-online learning
von: Sentenac, Flore, et al.
Veröffentlicht: (2025)
von: Sentenac, Flore, et al.
Veröffentlicht: (2025)
Federated Offline Reinforcement Learning: Collaborative Single-Policy Coverage Suffices
von: Woo, Jiin, et al.
Veröffentlicht: (2024)
von: Woo, Jiin, et al.
Veröffentlicht: (2024)
Ensemble sampling for linear bandits: small ensembles suffice
von: Janz, David, et al.
Veröffentlicht: (2023)
von: Janz, David, et al.
Veröffentlicht: (2023)
Optimistic Policy Optimization is Provably Efficient in Non-stationary MDPs
von: Zhong, Han, et al.
Veröffentlicht: (2021)
von: Zhong, Han, et al.
Veröffentlicht: (2021)
Linear Bellman Completeness Suffices for Efficient Online Reinforcement Learning with Few Actions
von: Golowich, Noah, et al.
Veröffentlicht: (2024)
von: Golowich, Noah, et al.
Veröffentlicht: (2024)
Inverse Q-Learning Done Right: Offline Imitation Learning in $Q^π$-Realizable MDPs
von: Moulin, Antoine, et al.
Veröffentlicht: (2025)
von: Moulin, Antoine, et al.
Veröffentlicht: (2025)
Exploration via linearly perturbed loss minimisation
von: Janz, David, et al.
Veröffentlicht: (2023)
von: Janz, David, et al.
Veröffentlicht: (2023)
Optimistic Actor-Critic with Parametric Policies for Linear Markov Decision Processes
von: Lin, Max Qiushi, et al.
Veröffentlicht: (2026)
von: Lin, Max Qiushi, et al.
Veröffentlicht: (2026)
Learning What to Recommend: Minimax Optimal Simple Regret in Logistic Bandits
von: Liu, Shuai, et al.
Veröffentlicht: (2026)
von: Liu, Shuai, et al.
Veröffentlicht: (2026)
Almost Free: Self-concordance in Natural Exponential Families and an Application to Bandits
von: Liu, Shuai, et al.
Veröffentlicht: (2024)
von: Liu, Shuai, et al.
Veröffentlicht: (2024)
Eluder dimension: localise it!
von: Bakhtiari, Alireza, et al.
Veröffentlicht: (2026)
von: Bakhtiari, Alireza, et al.
Veröffentlicht: (2026)
LACONIC: Length-Aware Constrained Reinforcement Learning for LLM
von: Liu, Chang, et al.
Veröffentlicht: (2026)
von: Liu, Chang, et al.
Veröffentlicht: (2026)
Querying Kernel Methods Suffices for Reconstructing their Training Data
von: Barzilai, Daniel, et al.
Veröffentlicht: (2025)
von: Barzilai, Daniel, et al.
Veröffentlicht: (2025)
Augmenting Offline RL with Unlabeled Data
von: Wang, Zhao, et al.
Veröffentlicht: (2024)
von: Wang, Zhao, et al.
Veröffentlicht: (2024)
Trajectory-Level Data Augmentation for Offline Reinforcement Learning
von: Schmähling, Tobias, et al.
Veröffentlicht: (2026)
von: Schmähling, Tobias, et al.
Veröffentlicht: (2026)
RoiRL: Efficient, Self-Supervised Reasoning with Offline Iterative Reinforcement Learning
von: Arzhantsev, Aleksei, et al.
Veröffentlicht: (2025)
von: Arzhantsev, Aleksei, et al.
Veröffentlicht: (2025)
To Believe or Not to Believe Your LLM
von: Yadkori, Yasin Abbasi, et al.
Veröffentlicht: (2024)
von: Yadkori, Yasin Abbasi, et al.
Veröffentlicht: (2024)
Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation
von: Lin, Max Qiushi, et al.
Veröffentlicht: (2025)
von: Lin, Max Qiushi, et al.
Veröffentlicht: (2025)
Offline Trajectory Optimization for Offline Reinforcement Learning
von: Zhao, Ziqi, et al.
Veröffentlicht: (2024)
von: Zhao, Ziqi, et al.
Veröffentlicht: (2024)
Subsampling Suffices for Adaptive Data Analysis
von: Blanc, Guy
Veröffentlicht: (2023)
von: Blanc, Guy
Veröffentlicht: (2023)
Stochastic Gradient Succeeds for Bandits
von: Mei, Jincheng, et al.
Veröffentlicht: (2024)
von: Mei, Jincheng, et al.
Veröffentlicht: (2024)
Ordering-based Conditions for Global Convergence of Policy Gradient Methods
von: Mei, Jincheng, et al.
Veröffentlicht: (2025)
von: Mei, Jincheng, et al.
Veröffentlicht: (2025)
Transformers on Markov Data: Constant Depth Suffices
von: Rajaraman, Nived, et al.
Veröffentlicht: (2024)
von: Rajaraman, Nived, et al.
Veröffentlicht: (2024)
Making Offline RL Online: Collaborative World Models for Offline Visual Reinforcement Learning
von: Wang, Qi, et al.
Veröffentlicht: (2023)
von: Wang, Qi, et al.
Veröffentlicht: (2023)
Logarithmic Width Suffices for Robust Memorization
von: Egosi, Amitsour, et al.
Veröffentlicht: (2025)
von: Egosi, Amitsour, et al.
Veröffentlicht: (2025)
Offline Inverse RL: New Solution Concepts and Provably Efficient Algorithms
von: Lazzati, Filippo, et al.
Veröffentlicht: (2024)
von: Lazzati, Filippo, et al.
Veröffentlicht: (2024)
Offline Reinforcement Learning: Role of State Aggregation and Trajectory Data
von: Jia, Zeyu, et al.
Veröffentlicht: (2024)
von: Jia, Zeyu, et al.
Veröffentlicht: (2024)
Accelerating Diffusion Planners in Offline RL via Reward-Aware Consistency Trajectory Distillation
von: Duan, Xintong, et al.
Veröffentlicht: (2025)
von: Duan, Xintong, et al.
Veröffentlicht: (2025)
Improving Offline RL by Blending Heuristics
von: Geng, Sinong, et al.
Veröffentlicht: (2023)
von: Geng, Sinong, et al.
Veröffentlicht: (2023)
Immediate Derivatives Suffice for Online Recurrent Adaptation
von: Merin, Aur Shalev
Veröffentlicht: (2026)
von: Merin, Aur Shalev
Veröffentlicht: (2026)
Ähnliche Einträge
-
Regret Minimization via Saddle Point Optimization
von: Kirschner, Johannes, et al.
Veröffentlicht: (2024) -
Confident Natural Policy Gradient for Local Planning in $q_π$-realizable Constrained MDPs
von: Tian, Tian, et al.
Veröffentlicht: (2024) -
Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence
von: György, András, et al.
Veröffentlicht: (2025) -
Rectifying Regression in Reinforcement Learning
von: Ayoub, Alex, et al.
Veröffentlicht: (2025) -
Sharper Guarantees for Misspecified Kernelized Bandit Optimization
von: Maran, Davide, et al.
Veröffentlicht: (2026)