MTA-RL: Robust Urban Driving via Multi-modal Transformer-based 3D Affordances and Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Guangli, Li, Dianzhao, Zhong, Wenjian, Xie, Bangquan, Okhrin, Ostap |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ethics-Aware Safe Reinforcement Learning for Rare-Event Risk Control in Interactive Urban Driving
by: Li, Dianzhao, et al.
Published: (2025)
by: Li, Dianzhao, et al.
Published: (2025)
A Platform-Agnostic Deep Reinforcement Learning Framework for Effective Sim2Real Transfer towards Autonomous Driving
by: Li, Dianzhao, et al.
Published: (2023)
by: Li, Dianzhao, et al.
Published: (2023)
Vision-based DRL Autonomous Driving Agent with Sim2Real Transfer
by: Li, Dianzhao, et al.
Published: (2023)
by: Li, Dianzhao, et al.
Published: (2023)
Autonomous Driving Small-Scale Cars: A Survey of Recent Development
by: Li, Dianzhao, et al.
Published: (2024)
by: Li, Dianzhao, et al.
Published: (2024)
Two-step dynamic obstacle avoidance
by: Hart, Fabian, et al.
Published: (2023)
by: Hart, Fabian, et al.
Published: (2023)
DAP: Diffusion-based Affordance Prediction for Multi-modality Storage
by: Chang, Haonan, et al.
Published: (2024)
by: Chang, Haonan, et al.
Published: (2024)
Affordance-R1: Reinforcement Learning for Generalizable Affordance Reasoning in Multimodal Large Language Model
by: Wang, Hanqing, et al.
Published: (2025)
by: Wang, Hanqing, et al.
Published: (2025)
Addressing Maximization Bias in Reinforcement Learning with Two-Sample Testing
by: Waltz, Martin, et al.
Published: (2022)
by: Waltz, Martin, et al.
Published: (2022)
AffordanceGrasp-R1:Leveraging Reasoning-Based Affordance Segmentation with Reinforcement Learning for Robotic Grasping
by: Zhou, Dingyi, et al.
Published: (2026)
by: Zhou, Dingyi, et al.
Published: (2026)
3D-AffordanceLLM: Harnessing Large Language Models for Open-Vocabulary Affordance Detection in 3D Worlds
by: Chu, Hengshuo, et al.
Published: (2025)
by: Chu, Hengshuo, et al.
Published: (2025)
MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning
by: Fu, Haoyu, et al.
Published: (2025)
by: Fu, Haoyu, et al.
Published: (2025)
DriveVLM-RL: Neuroscience-Inspired Reinforcement Learning with Vision-Language Models for Safe and Deployable Autonomous Driving
by: Huang, Zilin, et al.
Published: (2026)
by: Huang, Zilin, et al.
Published: (2026)
Distilling Multi-modal Large Language Models for Autonomous Driving
by: Hegde, Deepti, et al.
Published: (2025)
by: Hegde, Deepti, et al.
Published: (2025)
RAD: Training an End-to-End Driving Policy via Large-Scale 3DGS-based Reinforcement Learning
by: Gao, Hao, et al.
Published: (2025)
by: Gao, Hao, et al.
Published: (2025)
VoxAfford: Multi-Scale Voxel-Token Fusion for Open-Vocabulary 3D Affordance Detection
by: Sun, Haowen, et al.
Published: (2026)
by: Sun, Haowen, et al.
Published: (2026)
LM-MCVT: A Lightweight Multi-modal Multi-view Convolutional-Vision Transformer Approach for 3D Object Recognition
by: Xiong, Songsong, et al.
Published: (2025)
by: Xiong, Songsong, et al.
Published: (2025)
Driving Intents Amplify Planning-Oriented Reinforcement Learning
by: Lu, Hengtong, et al.
Published: (2026)
by: Lu, Hengtong, et al.
Published: (2026)
AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning
by: Jiang, Bo, et al.
Published: (2025)
by: Jiang, Bo, et al.
Published: (2025)
METDrive: Multi-modal End-to-end Autonomous Driving with Temporal Guidance
by: Guo, Ziang, et al.
Published: (2024)
by: Guo, Ziang, et al.
Published: (2024)
Interpretable Affordance Detection on 3D Point Clouds with Probabilistic Prototypes
by: Li, Maximilian Xiling, et al.
Published: (2025)
by: Li, Maximilian Xiling, et al.
Published: (2025)
2-Level Reinforcement Learning for Ships on Inland Waterways: Path Planning and Following
by: Waltz, Martin, et al.
Published: (2023)
by: Waltz, Martin, et al.
Published: (2023)
VLM-RL: A Unified Vision Language Models and Reinforcement Learning Framework for Safe Autonomous Driving
by: Huang, Zilin, et al.
Published: (2024)
by: Huang, Zilin, et al.
Published: (2024)
Affordance RAG: Hierarchical Multimodal Retrieval with Affordance-Aware Embodied Memory for Mobile Manipulation
by: Korekata, Ryosuke, et al.
Published: (2025)
by: Korekata, Ryosuke, et al.
Published: (2025)
Neural Rendering based Urban Scene Reconstruction for Autonomous Driving
by: Shen, Shihao, et al.
Published: (2024)
by: Shen, Shihao, et al.
Published: (2024)
MM-Gaussian: 3D Gaussian-based Multi-modal Fusion for Localization and Reconstruction in Unbounded Scenes
by: Wu, Chenyang, et al.
Published: (2024)
by: Wu, Chenyang, et al.
Published: (2024)
TRACER: Texture-Robust Affordance Chain-of-Thought for Deformable-Object Refinement
by: Jia, Wanjun, et al.
Published: (2026)
by: Jia, Wanjun, et al.
Published: (2026)
Learning Environment-Aware Affordance for 3D Articulated Object Manipulation under Occlusions
by: Wu, Ruihai, et al.
Published: (2023)
by: Wu, Ruihai, et al.
Published: (2023)
AffordanceLLM: Grounding Affordance from Vision Language Models
by: Qian, Shengyi, et al.
Published: (2024)
by: Qian, Shengyi, et al.
Published: (2024)
UniBEV: Multi-modal 3D Object Detection with Uniform BEV Encoders for Robustness against Missing Sensor Modalities
by: Wang, Shiming, et al.
Published: (2023)
by: Wang, Shiming, et al.
Published: (2023)
Learning-based 3D Reconstruction in Autonomous Driving: A Comprehensive Survey
by: Liao, Liewen, et al.
Published: (2025)
by: Liao, Liewen, et al.
Published: (2025)
Affordance-Guided Reinforcement Learning via Visual Prompting
by: Lee, Olivia Y., et al.
Published: (2024)
by: Lee, Olivia Y., et al.
Published: (2024)
Explanation for Trajectory Planning using Multi-modal Large Language Model for Autonomous Driving
by: Yamazaki, Shota, et al.
Published: (2024)
by: Yamazaki, Shota, et al.
Published: (2024)
CompassAD: Intent-Driven 3D Affordance Grounding in Functionally Competing Objects
by: Li, Jingliang, et al.
Published: (2026)
by: Li, Jingliang, et al.
Published: (2026)
Gradient-Driven 3D Segmentation and Affordance Transfer in Gaussian Splatting Using 2D Masks
by: Joseph, Joji, et al.
Published: (2024)
by: Joseph, Joji, et al.
Published: (2024)
Multi-modal Situated Reasoning in 3D Scenes
by: Linghu, Xiongkun, et al.
Published: (2024)
by: Linghu, Xiongkun, et al.
Published: (2024)
3D Object Visibility Prediction in Autonomous Driving
by: Luo, Chuanyu, et al.
Published: (2024)
by: Luo, Chuanyu, et al.
Published: (2024)
Multi-Objective Reinforcement Learning for Adaptable Personalized Autonomous Driving
by: Surmann, Hendrik, et al.
Published: (2025)
by: Surmann, Hendrik, et al.
Published: (2025)
Multi-modal Motion Prediction using Temporal Ensembling with Learning-based Aggregation
by: Hong, Kai-Yin, et al.
Published: (2024)
by: Hong, Kai-Yin, et al.
Published: (2024)
Panoramic Affordance Prediction
by: Zhang, Zixin, et al.
Published: (2026)
by: Zhang, Zixin, et al.
Published: (2026)
When LLMs step into the 3D World: A Survey and Meta-Analysis of 3D Tasks via Multi-modal Large Language Models
by: Ma, Xianzheng, et al.
Published: (2024)
by: Ma, Xianzheng, et al.
Published: (2024)
Similar Items
-
Ethics-Aware Safe Reinforcement Learning for Rare-Event Risk Control in Interactive Urban Driving
by: Li, Dianzhao, et al.
Published: (2025) -
A Platform-Agnostic Deep Reinforcement Learning Framework for Effective Sim2Real Transfer towards Autonomous Driving
by: Li, Dianzhao, et al.
Published: (2023) -
Vision-based DRL Autonomous Driving Agent with Sim2Real Transfer
by: Li, Dianzhao, et al.
Published: (2023) -
Autonomous Driving Small-Scale Cars: A Survey of Recent Development
by: Li, Dianzhao, et al.
Published: (2024) -
Two-step dynamic obstacle avoidance
by: Hart, Fabian, et al.
Published: (2023)