NeoRL: Efficient Exploration for Nonepisodic RL
Fuente:
arXiv
Guardado en:
| Autores principales: | Sukhija, Bhavya, Treven, Lenart, Dörfler, Florian, Coros, Stelian, Krause, Andreas |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Sample-efficient and Scalable Exploration in Continuous-Time RL
por: Iten, Klemens, et al.
Publicado: (2025)
por: Iten, Klemens, et al.
Publicado: (2025)
When to Sense and Control? A Time-adaptive Approach for Continuous-Time RL
por: Treven, Lenart, et al.
Publicado: (2024)
por: Treven, Lenart, et al.
Publicado: (2024)
Simulation Priors for Data-Efficient Deep Learning
por: Treven, Lenart, et al.
Publicado: (2025)
por: Treven, Lenart, et al.
Publicado: (2025)
Bridging the Sim-to-Real Gap with Bayesian Inference
por: Rothfuss, Jonas, et al.
Publicado: (2024)
por: Rothfuss, Jonas, et al.
Publicado: (2024)
ActSafe: Active Exploration with Safety Constraints for Reinforcement Learning
por: As, Yarden, et al.
Publicado: (2024)
por: As, Yarden, et al.
Publicado: (2024)
SOMBRL: Scalable and Optimistic Model-Based RL
por: Sukhija, Bhavya, et al.
Publicado: (2025)
por: Sukhija, Bhavya, et al.
Publicado: (2025)
TARC: Time-Adaptive Robotic Control
por: Sukhija, Arnav, et al.
Publicado: (2025)
por: Sukhija, Arnav, et al.
Publicado: (2025)
MaxInfoRL: Boosting exploration in reinforcement learning through information gain maximization
por: Sukhija, Bhavya, et al.
Publicado: (2024)
por: Sukhija, Bhavya, et al.
Publicado: (2024)
Active Few-Shot Fine-Tuning
por: Hübotter, Jonas, et al.
Publicado: (2024)
por: Hübotter, Jonas, et al.
Publicado: (2024)
Transductive Active Learning: Theory and Applications
por: Hübotter, Jonas, et al.
Publicado: (2024)
por: Hübotter, Jonas, et al.
Publicado: (2024)
Model-Based Reinforcement Learning for Control under Time-Varying Dynamics
por: Iten, Klemens, et al.
Publicado: (2026)
por: Iten, Klemens, et al.
Publicado: (2026)
Data-Efficient Task Generalization via Probabilistic Model-based Meta Reinforcement Learning
por: Bhardwaj, Arjun, et al.
Publicado: (2023)
por: Bhardwaj, Arjun, et al.
Publicado: (2023)
Safe Exploration Using Bayesian World Models and Log-Barrier Optimization
por: As, Yarden, et al.
Publicado: (2024)
por: As, Yarden, et al.
Publicado: (2024)
NeoRL-2: Near Real-World Benchmarks for Offline Reinforcement Learning with Extended Realistic Scenarios
por: Gao, Songyi, et al.
Publicado: (2025)
por: Gao, Songyi, et al.
Publicado: (2025)
Optimistic Online LQR via Intrinsic Rewards
por: Bartos, Marcell, et al.
Publicado: (2026)
por: Bartos, Marcell, et al.
Publicado: (2026)
Safe Exploration via Policy Priors
por: Wendl, Manuel, et al.
Publicado: (2026)
por: Wendl, Manuel, et al.
Publicado: (2026)
Maximum Entropy Behavior Exploration for Sim2Real Zero-Shot Reinforcement Learning
por: Hu, Jiajun, et al.
Publicado: (2026)
por: Hu, Jiajun, et al.
Publicado: (2026)
Problem Space Transformations for Out-of-Distribution Generalisation in Behavioural Cloning
por: Doshi, Kiran, et al.
Publicado: (2024)
por: Doshi, Kiran, et al.
Publicado: (2024)
Model-Based RL for Mean-Field Games is not Statistically Harder than Single-Agent RL
por: Huang, Jiawei, et al.
Publicado: (2024)
por: Huang, Jiawei, et al.
Publicado: (2024)
Learning More With Less: Sample Efficient Model-Based RL for Loco-Manipulation
por: Hoffman, Benjamin, et al.
Publicado: (2025)
por: Hoffman, Benjamin, et al.
Publicado: (2025)
Beyond Binary: Sim-to-Real Dexterous Manipulation with Physics-Grounded Contact Representation
por: Pan, Jiahe, et al.
Publicado: (2026)
por: Pan, Jiahe, et al.
Publicado: (2026)
CAIMAN: Causal Action Influence Detection for Sample-efficient Loco-manipulation
por: Yuan, Yuanchen, et al.
Publicado: (2025)
por: Yuan, Yuanchen, et al.
Publicado: (2025)
Epistemically-guided forward-backward exploration
por: Urpí, Núria Armengol, et al.
Publicado: (2025)
por: Urpí, Núria Armengol, et al.
Publicado: (2025)
MetaLoco: Universal Quadrupedal Locomotion with Meta-Reinforcement Learning and Motion Imitation
por: Zargarbashi, Fatemeh, et al.
Publicado: (2024)
por: Zargarbashi, Fatemeh, et al.
Publicado: (2024)
Neural Modes: Self-supervised Learning of Nonlinear Modal Subspaces
por: Wang, Jiahong, et al.
Publicado: (2024)
por: Wang, Jiahong, et al.
Publicado: (2024)
JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training
por: Hu, Zhengding, et al.
Publicado: (2026)
por: Hu, Zhengding, et al.
Publicado: (2026)
floq: Training Critics via Flow-Matching for Scaling Compute in Value-Based RL
por: Agrawalla, Bhavya, et al.
Publicado: (2025)
por: Agrawalla, Bhavya, et al.
Publicado: (2025)
Spectral Bellman Method: Unifying Representation and Exploration in RL
por: Nabati, Ofir, et al.
Publicado: (2025)
por: Nabati, Ofir, et al.
Publicado: (2025)
Meta-RL Induces Exploration in Language Agents
por: Jiang, Yulun, et al.
Publicado: (2025)
por: Jiang, Yulun, et al.
Publicado: (2025)
RobotKeyframing: Learning Locomotion with High-Level Objectives via Mixture of Dense and Sparse Rewards
por: Zargarbashi, Fatemeh, et al.
Publicado: (2024)
por: Zargarbashi, Fatemeh, et al.
Publicado: (2024)
Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models
por: Liu, Runze, et al.
Publicado: (2025)
por: Liu, Runze, et al.
Publicado: (2025)
Uncovering RL Integration in SSL Loss: Objective-Specific Implications for Data-Efficient RL
por: Çağatan, Ömer Veysel, et al.
Publicado: (2024)
por: Çağatan, Ömer Veysel, et al.
Publicado: (2024)
Demystifying the Mechanisms Behind Emergent Exploration in Goal-conditioned RL
por: Bastankhah, Mahsa, et al.
Publicado: (2025)
por: Bastankhah, Mahsa, et al.
Publicado: (2025)
Rethinking Robustness Assessment: Adversarial Attacks on Learning-based Quadrupedal Locomotion Controllers
por: Shi, Fan, et al.
Publicado: (2024)
por: Shi, Fan, et al.
Publicado: (2024)
Efficient RL Training for LLMs with Experience Replay
por: Arnal, Charles, et al.
Publicado: (2026)
por: Arnal, Charles, et al.
Publicado: (2026)
BroRL: Scaling Reinforcement Learning via Broadened Exploration
por: Hu, Jian, et al.
Publicado: (2025)
por: Hu, Jian, et al.
Publicado: (2025)
Exploration Hacking: Can LLMs Learn to Resist RL Training?
por: Jang, Eyon, et al.
Publicado: (2026)
por: Jang, Eyon, et al.
Publicado: (2026)
Improving RL Exploration for LLM Reasoning through Retrospective Replay
por: Dou, Shihan, et al.
Publicado: (2025)
por: Dou, Shihan, et al.
Publicado: (2025)
Token-Efficient RL for LLM Reasoning
por: Lee, Alan, et al.
Publicado: (2025)
por: Lee, Alan, et al.
Publicado: (2025)
RAMBO: RL-Augmented Model-Based Whole-Body Control for Loco-Manipulation
por: Cheng, Jin, et al.
Publicado: (2025)
por: Cheng, Jin, et al.
Publicado: (2025)
Ejemplares similares
-
Sample-efficient and Scalable Exploration in Continuous-Time RL
por: Iten, Klemens, et al.
Publicado: (2025) -
When to Sense and Control? A Time-adaptive Approach for Continuous-Time RL
por: Treven, Lenart, et al.
Publicado: (2024) -
Simulation Priors for Data-Efficient Deep Learning
por: Treven, Lenart, et al.
Publicado: (2025) -
Bridging the Sim-to-Real Gap with Bayesian Inference
por: Rothfuss, Jonas, et al.
Publicado: (2024) -
ActSafe: Active Exploration with Safety Constraints for Reinforcement Learning
por: As, Yarden, et al.
Publicado: (2024)