Eluder dimension: localise it!
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bakhtiari, Alireza, Ayoub, Alex, Robertson, Samuel, Janz, David, Szepesvári, Csaba |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Rectifying Regression in Reinforcement Learning
von: Ayoub, Alex, et al.
Veröffentlicht: (2025)
von: Ayoub, Alex, et al.
Veröffentlicht: (2025)
Learning What to Recommend: Minimax Optimal Simple Regret in Logistic Bandits
von: Liu, Shuai, et al.
Veröffentlicht: (2026)
von: Liu, Shuai, et al.
Veröffentlicht: (2026)
Exploration via linearly perturbed loss minimisation
von: Janz, David, et al.
Veröffentlicht: (2023)
von: Janz, David, et al.
Veröffentlicht: (2023)
Sharp analysis of linear ensemble sampling
von: Akhavan, Arya, et al.
Veröffentlicht: (2026)
von: Akhavan, Arya, et al.
Veröffentlicht: (2026)
Ensemble sampling for linear bandits: small ensembles suffice
von: Janz, David, et al.
Veröffentlicht: (2023)
von: Janz, David, et al.
Veröffentlicht: (2023)
Regret Minimization via Saddle Point Optimization
von: Kirschner, Johannes, et al.
Veröffentlicht: (2024)
von: Kirschner, Johannes, et al.
Veröffentlicht: (2024)
Learning to Reason Efficiently with Discounted Reinforcement Learning
von: Ayoub, Alex, et al.
Veröffentlicht: (2025)
von: Ayoub, Alex, et al.
Veröffentlicht: (2025)
Almost Free: Self-concordance in Natural Exponential Families and an Application to Bandits
von: Liu, Shuai, et al.
Veröffentlicht: (2024)
von: Liu, Shuai, et al.
Veröffentlicht: (2024)
Switching the Loss Reduces the Cost in Batch (Offline) Reinforcement Learning
von: Ayoub, Alex, et al.
Veröffentlicht: (2024)
von: Ayoub, Alex, et al.
Veröffentlicht: (2024)
Sharper Guarantees for Misspecified Kernelized Bandit Optimization
von: Maran, Davide, et al.
Veröffentlicht: (2026)
von: Maran, Davide, et al.
Veröffentlicht: (2026)
Eluder-based Regret for Stochastic Contextual MDPs
von: Levy, Orin, et al.
Veröffentlicht: (2022)
von: Levy, Orin, et al.
Veröffentlicht: (2022)
Trajectory Data Suffices for Statistically Efficient Learning in Offline RL with Linear $q^π$-Realizability and Concentrability
von: Tkachuk, Volodymyr, et al.
Veröffentlicht: (2024)
von: Tkachuk, Volodymyr, et al.
Veröffentlicht: (2024)
Balancing optimism and pessimism in offline-to-online learning
von: Sentenac, Flore, et al.
Veröffentlicht: (2025)
von: Sentenac, Flore, et al.
Veröffentlicht: (2025)
Confident Natural Policy Gradient for Local Planning in $q_π$-realizable Constrained MDPs
von: Tian, Tian, et al.
Veröffentlicht: (2024)
von: Tian, Tian, et al.
Veröffentlicht: (2024)
Stochastic Gradient Descent for Gaussian Processes Done Right
von: Lin, Jihao Andreas, et al.
Veröffentlicht: (2023)
von: Lin, Jihao Andreas, et al.
Veröffentlicht: (2023)
Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence
von: György, András, et al.
Veröffentlicht: (2025)
von: György, András, et al.
Veröffentlicht: (2025)
Optimistic Policy Optimization is Provably Efficient in Non-stationary MDPs
von: Zhong, Han, et al.
Veröffentlicht: (2021)
von: Zhong, Han, et al.
Veröffentlicht: (2021)
To Believe or Not to Believe Your LLM
von: Yadkori, Yasin Abbasi, et al.
Veröffentlicht: (2024)
von: Yadkori, Yasin Abbasi, et al.
Veröffentlicht: (2024)
Does Weighting Improve Matrix Factorization for Recommender Systems?
von: Ayoub, Alex, et al.
Veröffentlicht: (2025)
von: Ayoub, Alex, et al.
Veröffentlicht: (2025)
Ordering-based Conditions for Global Convergence of Policy Gradient Methods
von: Mei, Jincheng, et al.
Veröffentlicht: (2025)
von: Mei, Jincheng, et al.
Veröffentlicht: (2025)
Stochastic Gradient Succeeds for Bandits
von: Mei, Jincheng, et al.
Veröffentlicht: (2024)
von: Mei, Jincheng, et al.
Veröffentlicht: (2024)
LACONIC: Length-Aware Constrained Reinforcement Learning for LLM
von: Liu, Chang, et al.
Veröffentlicht: (2026)
von: Liu, Chang, et al.
Veröffentlicht: (2026)
Optimistic Actor-Critic with Parametric Policies for Linear Markov Decision Processes
von: Lin, Max Qiushi, et al.
Veröffentlicht: (2026)
von: Lin, Max Qiushi, et al.
Veröffentlicht: (2026)
Small steps no more: Global convergence of stochastic gradient bandits for arbitrary learning rates
von: Mei, Jincheng, et al.
Veröffentlicht: (2025)
von: Mei, Jincheng, et al.
Veröffentlicht: (2025)
Frontier LLMs Still Struggle with Simple Reasoning Tasks
von: Malek, Alan, et al.
Veröffentlicht: (2025)
von: Malek, Alan, et al.
Veröffentlicht: (2025)
Rethinking the Foundations for Continual Reinforcement Learning
von: Elelimy, Esraa, et al.
Veröffentlicht: (2025)
von: Elelimy, Esraa, et al.
Veröffentlicht: (2025)
Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation
von: Lin, Max Qiushi, et al.
Veröffentlicht: (2025)
von: Lin, Max Qiushi, et al.
Veröffentlicht: (2025)
Variance-sensitive Thompson sampling for generalised linear bandits, revisited
von: Perneczky, Tom, et al.
Veröffentlicht: (2026)
von: Perneczky, Tom, et al.
Veröffentlicht: (2026)
A Survey of State Representation Learning for Deep Reinforcement Learning
von: Echchahed, Ayoub, et al.
Veröffentlicht: (2025)
von: Echchahed, Ayoub, et al.
Veröffentlicht: (2025)
Bernstein-type dimension-free concentration for self-normalised martingales
von: Akhavan, Arya, et al.
Veröffentlicht: (2025)
von: Akhavan, Arya, et al.
Veröffentlicht: (2025)
When and why randomised exploration works (in linear bandits)
von: Abeille, Marc, et al.
Veröffentlicht: (2025)
von: Abeille, Marc, et al.
Veröffentlicht: (2025)
Revisiting Subgradient Dominance in Robust MDPs: Counterexamples, Hardness, and Sufficient Conditions
von: Kitamura, Toshinori, et al.
Veröffentlicht: (2026)
von: Kitamura, Toshinori, et al.
Veröffentlicht: (2026)
TV-SurvCaus: Dynamic Representation Balancing for Causal Survival Analysis
von: Abraich, Ayoub
Veröffentlicht: (2025)
von: Abraich, Ayoub
Veröffentlicht: (2025)
High-probability zeroth-order online convex optimisation beyond Euclidean geometry
von: Janz, David, et al.
Veröffentlicht: (2025)
von: Janz, David, et al.
Veröffentlicht: (2025)
Meta-models for transfer learning in source localisation
von: Bull, Lawrence A., et al.
Veröffentlicht: (2023)
von: Bull, Lawrence A., et al.
Veröffentlicht: (2023)
Scalable Machine Learning Algorithms using Path Signatures
von: Tóth, Csaba
Veröffentlicht: (2025)
von: Tóth, Csaba
Veröffentlicht: (2025)
Theoretical Guarantees for LT-TTD: A Unified Transformer-based Architecture for Two-Level Ranking Systems
von: Abraich, Ayoub
Veröffentlicht: (2025)
von: Abraich, Ayoub
Veröffentlicht: (2025)
Revisiting RIP guarantees for sketching operators on mixture models
von: Belhadji, Ayoub, et al.
Veröffentlicht: (2023)
von: Belhadji, Ayoub, et al.
Veröffentlicht: (2023)
Sketch and shift: a robust decoder for compressive clustering
von: Belhadji, Ayoub, et al.
Veröffentlicht: (2023)
von: Belhadji, Ayoub, et al.
Veröffentlicht: (2023)
Understanding Asynchronous Inference Methods for Vision-Language-Action Models
von: Agouzoul, Ayoub
Veröffentlicht: (2026)
von: Agouzoul, Ayoub
Veröffentlicht: (2026)
Ähnliche Einträge
-
Rectifying Regression in Reinforcement Learning
von: Ayoub, Alex, et al.
Veröffentlicht: (2025) -
Learning What to Recommend: Minimax Optimal Simple Regret in Logistic Bandits
von: Liu, Shuai, et al.
Veröffentlicht: (2026) -
Exploration via linearly perturbed loss minimisation
von: Janz, David, et al.
Veröffentlicht: (2023) -
Sharp analysis of linear ensemble sampling
von: Akhavan, Arya, et al.
Veröffentlicht: (2026) -
Ensemble sampling for linear bandits: small ensembles suffice
von: Janz, David, et al.
Veröffentlicht: (2023)