CaRL: Learning Scalable Planning Policies with Simple Rewards
Fuente:
arXiv
Guardado en:
| Autores principales: | Jaeger, Bernhard, Dauner, Daniel, Beißwenger, Jens, Gerstenecker, Simon, Chitta, Kashyap, Geiger, Andreas |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Hidden Biases of End-to-End Driving Datasets
por: Zimmerlin, Julian, et al.
Publicado: (2024)
por: Zimmerlin, Julian, et al.
Publicado: (2024)
SLEDGE: Synthesizing Driving Environments with Generative Models and Rule-Based Traffic
por: Chitta, Kashyap, et al.
Publicado: (2024)
por: Chitta, Kashyap, et al.
Publicado: (2024)
LEAD: Minimizing Learner-Expert Asymmetry in End-to-End Driving
por: Nguyen, Long, et al.
Publicado: (2025)
por: Nguyen, Long, et al.
Publicado: (2025)
End-to-end Autonomous Driving: Challenges and Frontiers
por: Chen, Li, et al.
Publicado: (2023)
por: Chen, Li, et al.
Publicado: (2023)
NAVSIM: Data-Driven Non-Reactive Autonomous Vehicle Simulation and Benchmarking
por: Dauner, Daniel, et al.
Publicado: (2024)
por: Dauner, Daniel, et al.
Publicado: (2024)
An Invitation to Deep Reinforcement Learning
por: Jaeger, Bernhard, et al.
Publicado: (2023)
por: Jaeger, Bernhard, et al.
Publicado: (2023)
Pseudo-Simulation for Autonomous Driving
por: Cao, Wei, et al.
Publicado: (2025)
por: Cao, Wei, et al.
Publicado: (2025)
PlanT 2.0: Exposing Biases and Structural Flaws in Closed-Loop Driving
por: Gerstenecker, Simon, et al.
Publicado: (2025)
por: Gerstenecker, Simon, et al.
Publicado: (2025)
Sample-efficient and Scalable Exploration in Continuous-Time RL
por: Iten, Klemens, et al.
Publicado: (2025)
por: Iten, Klemens, et al.
Publicado: (2025)
A Review of Online Diffusion Policy RL Algorithms for Scalable Robotic Control
por: Choi, Wonhyeok, et al.
Publicado: (2026)
por: Choi, Wonhyeok, et al.
Publicado: (2026)
Centaur: Robust End-to-End Autonomous Driving with Test-Time Training
por: Sima, Chonghao, et al.
Publicado: (2025)
por: Sima, Chonghao, et al.
Publicado: (2025)
Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making
por: Hori, Toshiaki, et al.
Publicado: (2025)
por: Hori, Toshiaki, et al.
Publicado: (2025)
Robot Policy Learning with Temporal Optimal Transport Reward
por: Fu, Yuwei, et al.
Publicado: (2024)
por: Fu, Yuwei, et al.
Publicado: (2024)
METRA: Scalable Unsupervised RL with Metric-Aware Abstraction
por: Park, Seohong, et al.
Publicado: (2023)
por: Park, Seohong, et al.
Publicado: (2023)
Fail2Drive: Benchmarking Closed-Loop Driving Generalization
por: Gerstenecker, Simon, et al.
Publicado: (2026)
por: Gerstenecker, Simon, et al.
Publicado: (2026)
SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
por: Li, Haozhan, et al.
Publicado: (2025)
por: Li, Haozhan, et al.
Publicado: (2025)
Towards Learning Scalable Agile Dynamic Motion Planning for Robosoccer Teams with Policy Optimization
por: Ho, Brandon, et al.
Publicado: (2025)
por: Ho, Brandon, et al.
Publicado: (2025)
ORSO: Accelerating Reward Design via Online Reward Selection and Policy Optimization
por: Zhang, Chen Bo Calvin, et al.
Publicado: (2024)
por: Zhang, Chen Bo Calvin, et al.
Publicado: (2024)
Confidence-Controlled Exploration: Efficient Sparse-Reward Policy Learning for Robot Navigation
por: Patel, Bhrij, et al.
Publicado: (2023)
por: Patel, Bhrij, et al.
Publicado: (2023)
DEAS: DEtached value learning with Action Sequence for Scalable Offline RL
por: Kim, Changyeon, et al.
Publicado: (2025)
por: Kim, Changyeon, et al.
Publicado: (2025)
DISCOVER: Automated Curricula for Sparse-Reward Reinforcement Learning
por: Diaz-Bone, Leander, et al.
Publicado: (2025)
por: Diaz-Bone, Leander, et al.
Publicado: (2025)
Posterior Behavioral Cloning: Pretraining BC Policies for Efficient RL Finetuning
por: Wagenmaker, Andrew, et al.
Publicado: (2025)
por: Wagenmaker, Andrew, et al.
Publicado: (2025)
Diffusion-Based Impedance Learning for Contact-Rich Manipulation Tasks
por: Geiger, Noah, et al.
Publicado: (2025)
por: Geiger, Noah, et al.
Publicado: (2025)
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only
por: Xiao, Wei, et al.
Publicado: (2025)
por: Xiao, Wei, et al.
Publicado: (2025)
Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions
por: Ishihara, Yu, et al.
Publicado: (2025)
por: Ishihara, Yu, et al.
Publicado: (2025)
Neural-Network-Driven Reward Prediction as a Heuristic: Advancing Q-Learning for Mobile Robot Path Planning
por: Ji, Yiming, et al.
Publicado: (2024)
por: Ji, Yiming, et al.
Publicado: (2024)
Reward Learning from Suboptimal Demonstrations with Applications in Surgical Electrocautery
por: Karimi, Zohre, et al.
Publicado: (2024)
por: Karimi, Zohre, et al.
Publicado: (2024)
Causally Robust Reward Learning from Reason-Augmented Preference Feedback
por: Hwang, Minjune, et al.
Publicado: (2026)
por: Hwang, Minjune, et al.
Publicado: (2026)
MASP: Scalable GNN-based Planning for Multi-Agent Navigation
por: Yang, Xinyi, et al.
Publicado: (2023)
por: Yang, Xinyi, et al.
Publicado: (2023)
Assuring the Safety of Reinforcement Learning Components: AMLAS-RL
por: Imrie, Calum Corrie, et al.
Publicado: (2025)
por: Imrie, Calum Corrie, et al.
Publicado: (2025)
Q-Guided Stein Variational Model Predictive Control via RL-informed Policy Prior
por: Cai, Shizhe, et al.
Publicado: (2025)
por: Cai, Shizhe, et al.
Publicado: (2025)
SPAARS: Safer RL Policy Alignment through Abstract Exploration and Refined Exploitation of Action Space
por: K, Swaminathan S, et al.
Publicado: (2026)
por: K, Swaminathan S, et al.
Publicado: (2026)
Diffusion-Reward Adversarial Imitation Learning
por: Lai, Chun-Mao, et al.
Publicado: (2024)
por: Lai, Chun-Mao, et al.
Publicado: (2024)
Discrete-Guided Diffusion for Scalable and Safe Multi-Robot Motion Planning
por: Liang, Jinhao, et al.
Publicado: (2025)
por: Liang, Jinhao, et al.
Publicado: (2025)
SPRINT: Scalable Policy Pre-Training via Language Instruction Relabeling
por: Zhang, Jesse, et al.
Publicado: (2023)
por: Zhang, Jesse, et al.
Publicado: (2023)
Align and Filter: Improving Performance in Asynchronous On-Policy RL
por: Honari, Homayoun, et al.
Publicado: (2026)
por: Honari, Homayoun, et al.
Publicado: (2026)
Reward-Punishment Reinforcement Learning with Maximum Entropy
por: Wang, Jiexin, et al.
Publicado: (2024)
por: Wang, Jiexin, et al.
Publicado: (2024)
TimeRewarder: Learning Dense Reward from Passive Videos via Frame-wise Temporal Distance
por: Liu, Yuyang, et al.
Publicado: (2025)
por: Liu, Yuyang, et al.
Publicado: (2025)
Beyond Scalar Rewards: Distributional Reinforcement Learning with Preordered Objectives for Safe and Reliable Autonomous Driving
por: Abouelazm, Ahmed, et al.
Publicado: (2026)
por: Abouelazm, Ahmed, et al.
Publicado: (2026)
Dual RL: Unification and New Methods for Reinforcement and Imitation Learning
por: Sikchi, Harshit, et al.
Publicado: (2023)
por: Sikchi, Harshit, et al.
Publicado: (2023)
Ejemplares similares
-
Hidden Biases of End-to-End Driving Datasets
por: Zimmerlin, Julian, et al.
Publicado: (2024) -
SLEDGE: Synthesizing Driving Environments with Generative Models and Rule-Based Traffic
por: Chitta, Kashyap, et al.
Publicado: (2024) -
LEAD: Minimizing Learner-Expert Asymmetry in End-to-End Driving
por: Nguyen, Long, et al.
Publicado: (2025) -
End-to-end Autonomous Driving: Challenges and Frontiers
por: Chen, Li, et al.
Publicado: (2023) -
NAVSIM: Data-Driven Non-Reactive Autonomous Vehicle Simulation and Benchmarking
por: Dauner, Daniel, et al.
Publicado: (2024)