GenRL: Multimodal-foundation world models for generalization in embodied agents
Fuente:
arXiv
Guardado en:
| Autores principales: | Mazzaglia, Pietro, Verbelen, Tim, Dhoedt, Bart, Courville, Aaron, Rajeswar, Sai |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Representing Positional Information in Generative World Models for Object Manipulation
por: Ferraro, Stefano, et al.
Publicado: (2024)
por: Ferraro, Stefano, et al.
Publicado: (2024)
Choreographer: Learning and Adapting Skills in Imagination
por: Mazzaglia, Pietro, et al.
Publicado: (2022)
por: Mazzaglia, Pietro, et al.
Publicado: (2022)
Contrastive Active Inference
por: Mazzaglia, Pietro, et al.
Publicado: (2021)
por: Mazzaglia, Pietro, et al.
Publicado: (2021)
Information-driven Affordance Discovery for Efficient Robotic Manipulation
por: Mazzaglia, Pietro, et al.
Publicado: (2024)
por: Mazzaglia, Pietro, et al.
Publicado: (2024)
Learning Dynamic Cognitive Map with Autonomous Navigation
por: de Tinguy, Daria, et al.
Publicado: (2024)
por: de Tinguy, Daria, et al.
Publicado: (2024)
Exploring and Learning Structure: Active Inference Approach in Navigational Agents
por: de Tinguy, Daria, et al.
Publicado: (2024)
por: de Tinguy, Daria, et al.
Publicado: (2024)
Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models
por: Bendikas, Rokas, et al.
Publicado: (2025)
por: Bendikas, Rokas, et al.
Publicado: (2025)
Redundancy-aware Action Spaces for Robot Learning
por: Mazzaglia, Pietro, et al.
Publicado: (2024)
por: Mazzaglia, Pietro, et al.
Publicado: (2024)
Hybrid Training for Vision-Language-Action Models
por: Mazzaglia, Pietro, et al.
Publicado: (2025)
por: Mazzaglia, Pietro, et al.
Publicado: (2025)
Navigation and Exploration with Active Inference: from Biology to Industry
por: de Tinguy, Daria, et al.
Publicado: (2025)
por: de Tinguy, Daria, et al.
Publicado: (2025)
Towards foundational LiDAR world models with efficient latent flow matching
por: Liu, Tianran, et al.
Publicado: (2025)
por: Liu, Tianran, et al.
Publicado: (2025)
Virtual avatar generation models as world navigators
por: Mandava, Sai
Publicado: (2024)
por: Mandava, Sai
Publicado: (2024)
Thinker: A vision-language foundation model for embodied intelligence
por: Pan, Baiyu, et al.
Publicado: (2026)
por: Pan, Baiyu, et al.
Publicado: (2026)
Bio-Inspired Topological Autonomous Navigation with Active Inference in Robotics
por: de Tinguy, Daria, et al.
Publicado: (2025)
por: de Tinguy, Daria, et al.
Publicado: (2025)
Online Structure Learning and Planning for Autonomous Robot Navigation using Active Inference
por: de tinguy, Daria, et al.
Publicado: (2025)
por: de tinguy, Daria, et al.
Publicado: (2025)
OPFormer: Object Pose Estimation leveraging foundation model with geometric encoding
por: Moroz, Artem, et al.
Publicado: (2025)
por: Moroz, Artem, et al.
Publicado: (2025)
GenFlowRL: Shaping Rewards with Generative Object-Centric Flow in Visual Reinforcement Learning
por: Yu, Kelin, et al.
Publicado: (2025)
por: Yu, Kelin, et al.
Publicado: (2025)
Integrating Deep RL and Bayesian Inference for ObjectNav in Mobile Robotics
por: Castelo-Branco, João, et al.
Publicado: (2026)
por: Castelo-Branco, João, et al.
Publicado: (2026)
Empowering Embodied Visual Tracking with Visual Foundation Models and Offline RL
por: Zhong, Fangwei, et al.
Publicado: (2024)
por: Zhong, Fangwei, et al.
Publicado: (2024)
Validation & Exploration of Multimodal Deep-Learning Camera-Lidar Calibration models
por: Karramreddy, Venkat, et al.
Publicado: (2024)
por: Karramreddy, Venkat, et al.
Publicado: (2024)
DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous Driving
por: Zhou, Yang, et al.
Publicado: (2026)
por: Zhou, Yang, et al.
Publicado: (2026)
GenNBV: Generalizable Next-Best-View Policy for Active 3D Reconstruction
por: Chen, Xiao, et al.
Publicado: (2024)
por: Chen, Xiao, et al.
Publicado: (2024)
Part-Guided 3D RL for Sim2Real Articulated Object Manipulation
por: Xie, Pengwei, et al.
Publicado: (2024)
por: Xie, Pengwei, et al.
Publicado: (2024)
ActiveVLN: Towards Active Exploration via Multi-Turn RL in Vision-and-Language Navigation
por: Zhang, Zekai, et al.
Publicado: (2025)
por: Zhang, Zekai, et al.
Publicado: (2025)
VLM-RL: A Unified Vision Language Models and Reinforcement Learning Framework for Safe Autonomous Driving
por: Huang, Zilin, et al.
Publicado: (2024)
por: Huang, Zilin, et al.
Publicado: (2024)
DriveVLM-RL: Neuroscience-Inspired Reinforcement Learning with Vision-Language Models for Safe and Deployable Autonomous Driving
por: Huang, Zilin, et al.
Publicado: (2026)
por: Huang, Zilin, et al.
Publicado: (2026)
MTA-RL: Robust Urban Driving via Multi-modal Transformer-based 3D Affordances and Reinforcement Learning
por: Chen, Guangli, et al.
Publicado: (2026)
por: Chen, Guangli, et al.
Publicado: (2026)
SurgTrack: CAD-Free 3D Tracking of Real-world Surgical Instruments
por: Guo, Wenwu, et al.
Publicado: (2024)
por: Guo, Wenwu, et al.
Publicado: (2024)
Spatial and Temporal Hierarchy for Autonomous Navigation using Active Inference in Minigrid Environment
por: de Tinguy, Daria, et al.
Publicado: (2023)
por: de Tinguy, Daria, et al.
Publicado: (2023)
Surg-InvNeRF: Invertible NeRF for 3D tracking and reconstruction in surgical vision
por: Loza, Gerardo, et al.
Publicado: (2025)
por: Loza, Gerardo, et al.
Publicado: (2025)
GraspClutter6D: A Large-scale Real-world Dataset for Robust Perception and Grasping in Cluttered Scenes
por: Back, Seunghyeok, et al.
Publicado: (2025)
por: Back, Seunghyeok, et al.
Publicado: (2025)
Toward Aligning Human and Robot Actions via Multi-Modal Demonstration Learning
por: Zahid, Azizul, et al.
Publicado: (2025)
por: Zahid, Azizul, et al.
Publicado: (2025)
Articulated 3D Scene Graphs for Open-World Mobile Manipulation
por: Büchner, Martin, et al.
Publicado: (2026)
por: Büchner, Martin, et al.
Publicado: (2026)
MARS: Multimodal Active Robotic Sensing for Articulated Characterization
por: Zeng, Hongliang, et al.
Publicado: (2024)
por: Zeng, Hongliang, et al.
Publicado: (2024)
UNIC: Learning Unified Multimodal Extrinsic Contact Estimation
por: Xu, Zhengtong, et al.
Publicado: (2026)
por: Xu, Zhengtong, et al.
Publicado: (2026)
3D Scene Rendering with Multimodal Gaussian Splatting
por: Gau, Chi-Shiang, et al.
Publicado: (2026)
por: Gau, Chi-Shiang, et al.
Publicado: (2026)
Diffusion Models as Optimizers for Efficient Planning in Offline RL
por: Huang, Renming, et al.
Publicado: (2024)
por: Huang, Renming, et al.
Publicado: (2024)
Enhancing Vision-Language Models with Scene Graphs for Traffic Accident Understanding
por: Lohner, Aaron, et al.
Publicado: (2024)
por: Lohner, Aaron, et al.
Publicado: (2024)
LIAM: Multimodal Transformer for Language Instructions, Images, Actions and Semantic Maps
por: Wang, Yihao, et al.
Publicado: (2025)
por: Wang, Yihao, et al.
Publicado: (2025)
Multimodal Object Detection using Depth and Image Data for Manufacturing Parts
por: Mahjourian, Nazanin, et al.
Publicado: (2024)
por: Mahjourian, Nazanin, et al.
Publicado: (2024)
Ejemplares similares
-
Representing Positional Information in Generative World Models for Object Manipulation
por: Ferraro, Stefano, et al.
Publicado: (2024) -
Choreographer: Learning and Adapting Skills in Imagination
por: Mazzaglia, Pietro, et al.
Publicado: (2022) -
Contrastive Active Inference
por: Mazzaglia, Pietro, et al.
Publicado: (2021) -
Information-driven Affordance Discovery for Efficient Robotic Manipulation
por: Mazzaglia, Pietro, et al.
Publicado: (2024) -
Learning Dynamic Cognitive Map with Autonomous Navigation
por: de Tinguy, Daria, et al.
Publicado: (2024)