Game On: Towards Language Models as RL Experimenters
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Jingwei, Lampe, Thomas, Abdolmaleki, Abbas, Springenberg, Jost Tobias, Riedmiller, Martin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Offline Actor-Critic Reinforcement Learning Scales to Large Models
por: Springenberg, Jost Tobias, et al.
Publicado: (2024)
por: Springenberg, Jost Tobias, et al.
Publicado: (2024)
RL Token: Bootstrapping Online RL with Vision-Language-Action Models
por: Xu, Charles, et al.
Publicado: (2026)
por: Xu, Charles, et al.
Publicado: (2026)
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)
por: Qin, Chongli, et al.
Publicado: (2025)
por: Qin, Chongli, et al.
Publicado: (2025)
Less is more -- the Dispatcher/ Executor principle for multi-task Reinforcement Learning
por: Riedmiller, Martin, et al.
Publicado: (2023)
por: Riedmiller, Martin, et al.
Publicado: (2023)
Real-World Fluid Directed Rigid Body Control via Deep Reinforcement Learning
por: Bhardwaj, Mohak, et al.
Publicado: (2024)
por: Bhardwaj, Mohak, et al.
Publicado: (2024)
GATS: Gather-Attend-Scatter
por: Zolna, Konrad, et al.
Publicado: (2024)
por: Zolna, Konrad, et al.
Publicado: (2024)
VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning
por: Lu, Guanxing, et al.
Publicado: (2025)
por: Lu, Guanxing, et al.
Publicado: (2025)
HighwayLLM: Decision-Making and Navigation in Highway Driving with RL-Informed Language Model
por: Yildirim, Mustafa, et al.
Publicado: (2024)
por: Yildirim, Mustafa, et al.
Publicado: (2024)
ResWM: Residual-Action World Model for Visual RL
por: Zhang, Jseen, et al.
Publicado: (2026)
por: Zhang, Jseen, et al.
Publicado: (2026)
Discover, Learn, and Reinforce: Scaling Vision-Language-Action Pretraining with Diverse RL-Generated Trajectories
por: Yang, Rushuai, et al.
Publicado: (2025)
por: Yang, Rushuai, et al.
Publicado: (2025)
GameVLM: A Decision-making Framework for Robotic Task Planning Based on Visual Language Models and Zero-sum Games
por: Mei, Aoran, et al.
Publicado: (2024)
por: Mei, Aoran, et al.
Publicado: (2024)
RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback
por: Wang, Yufei, et al.
Publicado: (2024)
por: Wang, Yufei, et al.
Publicado: (2024)
Coupled Local and Global World Models for Efficient First Order RL
por: Amigo, Joseph, et al.
Publicado: (2026)
por: Amigo, Joseph, et al.
Publicado: (2026)
WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL
por: Jiang, Zhennan, et al.
Publicado: (2026)
por: Jiang, Zhennan, et al.
Publicado: (2026)
Zonal RL-RRT: Integrated RL-RRT Path Planning with Collision Probability and Zone Connectivity
por: Tahmasbi, AmirMohammad, et al.
Publicado: (2024)
por: Tahmasbi, AmirMohammad, et al.
Publicado: (2024)
SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning
por: Zhang, Borong, et al.
Publicado: (2025)
por: Zhang, Borong, et al.
Publicado: (2025)
ActiveVLN: Towards Active Exploration via Multi-Turn RL in Vision-and-Language Navigation
por: Zhang, Zekai, et al.
Publicado: (2025)
por: Zhang, Zekai, et al.
Publicado: (2025)
Trajectory Entropy: Modeling Game State Stability from Multimodality Trajectory Prediction
por: Zhang, Yesheng, et al.
Publicado: (2025)
por: Zhang, Yesheng, et al.
Publicado: (2025)
Investigating Memory in Model-Free RL with POPGym Arcade
por: Wang, Zekang, et al.
Publicado: (2025)
por: Wang, Zekang, et al.
Publicado: (2025)
Language-Conditioned Offline RL for Multi-Robot Navigation
por: Morad, Steven, et al.
Publicado: (2024)
por: Morad, Steven, et al.
Publicado: (2024)
Towards Backdoor-Based Ownership Verification for Vision-Language-Action Models
por: Sun, Ming, et al.
Publicado: (2026)
por: Sun, Ming, et al.
Publicado: (2026)
Embodied AI in Mobile Robots: Coverage Path Planning with Large Language Models
por: Kong, Xiangrui, et al.
Publicado: (2024)
por: Kong, Xiangrui, et al.
Publicado: (2024)
World4RL: Diffusion World Models for Policy Refinement with Reinforcement Learning for Robotic Manipulation
por: Jiang, Zhennan, et al.
Publicado: (2025)
por: Jiang, Zhennan, et al.
Publicado: (2025)
Towards Interpretable Visuo-Tactile Predictive Models for Soft Robot Interactions
por: Donato, Enrico, et al.
Publicado: (2024)
por: Donato, Enrico, et al.
Publicado: (2024)
Learning Robot Soccer from Egocentric Vision with Deep Reinforcement Learning
por: Tirumala, Dhruva, et al.
Publicado: (2024)
por: Tirumala, Dhruva, et al.
Publicado: (2024)
Towards Consistent and Explainable Motion Prediction using Heterogeneous Graph Attention
por: Demmler, Tobias, et al.
Publicado: (2024)
por: Demmler, Tobias, et al.
Publicado: (2024)
SLAC: Simulation-Pretrained Latent Action Space for Whole-Body Real-World RL
por: Hu, Jiaheng, et al.
Publicado: (2025)
por: Hu, Jiaheng, et al.
Publicado: (2025)
OffRIPP: Offline RL-based Informative Path Planning
por: Gadipudi, Srikar Babu, et al.
Publicado: (2024)
por: Gadipudi, Srikar Babu, et al.
Publicado: (2024)
Depth-Constrained ASV Navigation with Deep RL and Limited Sensing
por: Zhalehmehrabi, Amirhossein, et al.
Publicado: (2025)
por: Zhalehmehrabi, Amirhossein, et al.
Publicado: (2025)
UniT: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World Modeling
por: Chen, Boyu, et al.
Publicado: (2026)
por: Chen, Boyu, et al.
Publicado: (2026)
Never too Prim to Swim: An LLM-Enhanced RL-based Adaptive S-Surface Controller for AUVs under Extreme Sea Conditions
por: Xie, Guanwen, et al.
Publicado: (2025)
por: Xie, Guanwen, et al.
Publicado: (2025)
BUMBLE: Unifying Reasoning and Acting with Vision-Language Models for Building-wide Mobile Manipulation
por: Shah, Rutav, et al.
Publicado: (2024)
por: Shah, Rutav, et al.
Publicado: (2024)
KineVLA: Towards Kinematics-Aware Vision-Language-Action Models with Bi-Level Action Decomposition
por: Han, Gaoge, et al.
Publicado: (2026)
por: Han, Gaoge, et al.
Publicado: (2026)
On Time-Indexing as Inductive Bias in Deep RL for Sequential Manipulation Tasks
por: Qureshi, M. Nomaan, et al.
Publicado: (2024)
por: Qureshi, M. Nomaan, et al.
Publicado: (2024)
Online Behavior Modification for Expressive User Control of RL-Trained Robots
por: Sheidlower, Isaac, et al.
Publicado: (2024)
por: Sheidlower, Isaac, et al.
Publicado: (2024)
Large-Language-Model-Guided State Estimation for Partially Observable Task and Motion Planning
por: Kim, Yoonwoo, et al.
Publicado: (2026)
por: Kim, Yoonwoo, et al.
Publicado: (2026)
Towards Interactive and Learnable Cooperative Driving Automation: a Large Language Model-Driven Decision-Making Framework
por: Fang, Shiyu, et al.
Publicado: (2024)
por: Fang, Shiyu, et al.
Publicado: (2024)
HiFi-CS: Towards Open Vocabulary Visual Grounding For Robotic Grasping Using Vision-Language Models
por: Bhat, Vineet, et al.
Publicado: (2024)
por: Bhat, Vineet, et al.
Publicado: (2024)
First Order Model-Based RL through Decoupled Backpropagation
por: Amigo, Joseph, et al.
Publicado: (2025)
por: Amigo, Joseph, et al.
Publicado: (2025)
Casper: Inferring Diverse Intents for Assistive Teleoperation with Vision Language Models
por: Liu, Huihan, et al.
Publicado: (2025)
por: Liu, Huihan, et al.
Publicado: (2025)
Ejemplares similares
-
Offline Actor-Critic Reinforcement Learning Scales to Large Models
por: Springenberg, Jost Tobias, et al.
Publicado: (2024) -
RL Token: Bootstrapping Online RL with Vision-Language-Action Models
por: Xu, Charles, et al.
Publicado: (2026) -
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)
por: Qin, Chongli, et al.
Publicado: (2025) -
Less is more -- the Dispatcher/ Executor principle for multi-task Reinforcement Learning
por: Riedmiller, Martin, et al.
Publicado: (2023) -
Real-World Fluid Directed Rigid Body Control via Deep Reinforcement Learning
por: Bhardwaj, Mohak, et al.
Publicado: (2024)