Linear Bellman Completeness Suffices for Efficient Online Reinforcement Learning with Few Actions
Fuente:
arXiv
Guardado en:
| Autores principales: | Golowich, Noah, Moitra, Ankur |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The Role of Inherent Bellman Error in Offline Reinforcement Learning with Linear Function Approximation
por: Golowich, Noah, et al.
Publicado: (2024)
por: Golowich, Noah, et al.
Publicado: (2024)
Edit Distance Robust Watermarks via Indexing Pseudorandom Codes
por: Golowich, Noah, et al.
Publicado: (2024)
por: Golowich, Noah, et al.
Publicado: (2024)
Subliminal Effects in Your Data: A General Mechanism via Log-Linearity
por: Aden-Ali, Ishaq, et al.
Publicado: (2026)
por: Aden-Ali, Ishaq, et al.
Publicado: (2026)
Exploration is Harder than Prediction: Cryptographically Separating Reinforcement Learning from Supervised Learning
por: Golowich, Noah, et al.
Publicado: (2024)
por: Golowich, Noah, et al.
Publicado: (2024)
Is Efficient PAC Learning Possible with an Oracle That Responds 'Yes' or 'No'?
por: Daskalakis, Constantinos, et al.
Publicado: (2024)
por: Daskalakis, Constantinos, et al.
Publicado: (2024)
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning
por: Omura, Motoki, et al.
Publicado: (2025)
por: Omura, Motoki, et al.
Publicado: (2025)
From External to Swap Regret 2.0: An Efficient Reduction and Oblivious Adversary for Large Action Spaces
por: Dagan, Yuval, et al.
Publicado: (2023)
por: Dagan, Yuval, et al.
Publicado: (2023)
Symmetric Q-learning: Reducing Skewness of Bellman Error in Online Reinforcement Learning
por: Omura, Motoki, et al.
Publicado: (2024)
por: Omura, Motoki, et al.
Publicado: (2024)
Model Stealing for Any Low-Rank Language Model
por: Liu, Allen, et al.
Publicado: (2024)
por: Liu, Allen, et al.
Publicado: (2024)
Provably Learning from Modern Language Models via Low Logit Rank
por: Golowich, Noah, et al.
Publicado: (2025)
por: Golowich, Noah, et al.
Publicado: (2025)
Theoretical Barriers in Bellman-Based Reinforcement Learning
por: Pinon, Brieuc, et al.
Publicado: (2025)
por: Pinon, Brieuc, et al.
Publicado: (2025)
Sequences of Logits Reveal the Low Rank Structure of Language Models
por: Golowich, Noah, et al.
Publicado: (2025)
por: Golowich, Noah, et al.
Publicado: (2025)
Bellman Unbiasedness: Toward Provably Efficient Distributional Reinforcement Learning with General Value Function Approximation
por: Cho, Taehyun, et al.
Publicado: (2024)
por: Cho, Taehyun, et al.
Publicado: (2024)
The Hidden Game Problem
por: Buzaglo, Gon, et al.
Publicado: (2025)
por: Buzaglo, Gon, et al.
Publicado: (2025)
Near-Optimal Learning and Planning in Separated Latent MDPs
por: Chen, Fan, et al.
Publicado: (2024)
por: Chen, Fan, et al.
Publicado: (2024)
RLSynC: Offline-Online Reinforcement Learning for Synthon Completion
por: Baker, Frazier N., et al.
Publicado: (2023)
por: Baker, Frazier N., et al.
Publicado: (2023)
LLM-Generated Explanations Do Not Suffice for Ultra-Strong Machine Learning
por: Ai, Lun, et al.
Publicado: (2025)
por: Ai, Lun, et al.
Publicado: (2025)
MICRO: Model-Based Offline Reinforcement Learning with a Conservative Bellman Operator
por: Liu, Xiao-Yin, et al.
Publicado: (2023)
por: Liu, Xiao-Yin, et al.
Publicado: (2023)
Iterated $Q$-Network: Beyond One-Step Bellman Updates in Deep Reinforcement Learning
por: Vincent, Théo, et al.
Publicado: (2024)
por: Vincent, Théo, et al.
Publicado: (2024)
A Generalized Projected Bellman Error for Off-policy Value Estimation in Reinforcement Learning
por: Patterson, Andrew, et al.
Publicado: (2021)
por: Patterson, Andrew, et al.
Publicado: (2021)
Path-Coupled Bellman Flows for Distributional Reinforcement Learning
por: Xu, Boyang, et al.
Publicado: (2026)
por: Xu, Boyang, et al.
Publicado: (2026)
The Multiple Ticket Hypothesis: Random Sparse Subnetworks Suffice for RLVR
por: Adewuyi, Israel, et al.
Publicado: (2026)
por: Adewuyi, Israel, et al.
Publicado: (2026)
AGaLiTe: Approximate Gated Linear Transformers for Online Reinforcement Learning
por: Pramanik, Subhojeet, et al.
Publicado: (2023)
por: Pramanik, Subhojeet, et al.
Publicado: (2023)
Bellman Error Centering
por: Chen, Xingguo, et al.
Publicado: (2025)
por: Chen, Xingguo, et al.
Publicado: (2025)
Teaching RL Agents to Act Better: VLM as Action Advisor for Online Reinforcement Learning
por: Wu, Xiefeng, et al.
Publicado: (2025)
por: Wu, Xiefeng, et al.
Publicado: (2025)
The Role of Sparsity for Length Generalization in Transformers
por: Golowich, Noah, et al.
Publicado: (2025)
por: Golowich, Noah, et al.
Publicado: (2025)
Provably Efficient Action-Manipulation Attack Against Continuous Reinforcement Learning
por: Luo, Zhi, et al.
Publicado: (2024)
por: Luo, Zhi, et al.
Publicado: (2024)
SAMG: Offline-to-Online Reinforcement Learning via State-Action-Conditional Offline Model Guidance
por: Zhang, Liyu, et al.
Publicado: (2024)
por: Zhang, Liyu, et al.
Publicado: (2024)
LoRA-One: One-Step Full Gradient Could Suffice for Fine-Tuning Large Language Models, Provably and Efficiently
por: Zhang, Yuanhe, et al.
Publicado: (2025)
por: Zhang, Yuanhe, et al.
Publicado: (2025)
Parameterized Projected Bellman Operator
por: Vincent, Théo, et al.
Publicado: (2023)
por: Vincent, Théo, et al.
Publicado: (2023)
Non-Linear Reinforcement Learning in Large Action Spaces: Structural Conditions and Sample-efficiency of Posterior Sampling
por: Agarwal, Alekh, et al.
Publicado: (2022)
por: Agarwal, Alekh, et al.
Publicado: (2022)
On Learning Parities with Dependent Noise
por: Golowich, Noah, et al.
Publicado: (2024)
por: Golowich, Noah, et al.
Publicado: (2024)
Learning Rate Matters: Vanilla LoRA May Suffice for LLM Fine-tuning
por: Lee, Yu-Ang, et al.
Publicado: (2026)
por: Lee, Yu-Ang, et al.
Publicado: (2026)
Sample and Oracle Efficient Reinforcement Learning for MDPs with Linearly-Realizable Value Functions
por: Mhammedi, Zakaria
Publicado: (2024)
por: Mhammedi, Zakaria
Publicado: (2024)
Language-Guided Reinforcement Learning for Hard Attention in Few-Shot Learning
por: Nikpour, Bahareh, et al.
Publicado: (2023)
por: Nikpour, Bahareh, et al.
Publicado: (2023)
Bridging the Sim-to-Real Gap in Reinforcement Learning-Based Industrial Dispatching through Execution Semantics
por: Hoss, Jonathan, et al.
Publicado: (2026)
por: Hoss, Jonathan, et al.
Publicado: (2026)
Action-List Reinforcement Learning Syndrome Decoding for Binary Linear Block Codes
por: Taghipour, Milad, et al.
Publicado: (2025)
por: Taghipour, Milad, et al.
Publicado: (2025)
Towards Efficient Online Tuning of VLM Agents via Counterfactual Soft Reinforcement Learning
por: Feng, Lang, et al.
Publicado: (2025)
por: Feng, Lang, et al.
Publicado: (2025)
Reinforcement Learning with Action Chunking
por: Li, Qiyang, et al.
Publicado: (2025)
por: Li, Qiyang, et al.
Publicado: (2025)
The Power of Resets in Online Reinforcement Learning
por: Mhammedi, Zakaria, et al.
Publicado: (2024)
por: Mhammedi, Zakaria, et al.
Publicado: (2024)
Ejemplares similares
-
The Role of Inherent Bellman Error in Offline Reinforcement Learning with Linear Function Approximation
por: Golowich, Noah, et al.
Publicado: (2024) -
Edit Distance Robust Watermarks via Indexing Pseudorandom Codes
por: Golowich, Noah, et al.
Publicado: (2024) -
Subliminal Effects in Your Data: A General Mechanism via Log-Linearity
por: Aden-Ali, Ishaq, et al.
Publicado: (2026) -
Exploration is Harder than Prediction: Cryptographically Separating Reinforcement Learning from Supervised Learning
por: Golowich, Noah, et al.
Publicado: (2024) -
Is Efficient PAC Learning Possible with an Oracle That Responds 'Yes' or 'No'?
por: Daskalakis, Constantinos, et al.
Publicado: (2024)