Is a 3D-Tokenized LLM the Key to Reliable Autonomous Driving?
Fuente:
arXiv
Saved in:
| Main Authors: | Bai, Yifan, Wu, Dongming, Liu, Yingfei, Jia, Fan, Mao, Weixin, Zhang, Ziheng, Zhao, Yucheng, Shen, Jianbing, Wei, Xing, Wang, Tiancai, Zhang, Xiangyu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PADriver: Towards Personalized Autonomous Driving
by: Kou, Genghua, et al.
Published: (2025)
by: Kou, Genghua, et al.
Published: (2025)
Language Prompt for Autonomous Driving
by: Wu, Dongming, et al.
Published: (2023)
by: Wu, Dongming, et al.
Published: (2023)
SubjectDrive: Scaling Generative Data in Autonomous Driving via Subject Control
by: Huang, Binyuan, et al.
Published: (2024)
by: Huang, Binyuan, et al.
Published: (2024)
Panacea+: Panoramic and Controllable Video Generation for Autonomous Driving
by: Wen, Yuqing, et al.
Published: (2024)
by: Wen, Yuqing, et al.
Published: (2024)
Glad: A Streaming Scene Generator for Autonomous Driving
by: Xie, Bin, et al.
Published: (2025)
by: Xie, Bin, et al.
Published: (2025)
Stream Query Denoising for Vectorized HD Map Construction
by: Wang, Shuo, et al.
Published: (2024)
by: Wang, Shuo, et al.
Published: (2024)
Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness
by: Wang, Haochen, et al.
Published: (2025)
by: Wang, Haochen, et al.
Published: (2025)
DEGround: An Effective Baseline for Ego-centric 3D Visual Grounding with a Homogeneous Framework
by: Zhang, Yani, et al.
Published: (2025)
by: Zhang, Yani, et al.
Published: (2025)
Hita: Holistic Tokenizer for Autoregressive Image Generation
by: Zheng, Anlin, et al.
Published: (2025)
by: Zheng, Anlin, et al.
Published: (2025)
RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping
by: Wu, Dongming, et al.
Published: (2025)
by: Wu, Dongming, et al.
Published: (2025)
Accelerating Training of Autoregressive Video Generation Models via Local Optimization with Representation Continuity
by: Zhou, Yucheng, et al.
Published: (2026)
by: Zhou, Yucheng, et al.
Published: (2026)
DME-Driver: Integrating Human Decision Logic and 3D Scene Perception in Autonomous Driving
by: Han, Wencheng, et al.
Published: (2024)
by: Han, Wencheng, et al.
Published: (2024)
SpatialActor: Exploring Disentangled Spatial Representations for Robust Robotic Manipulation
by: Shi, Hao, et al.
Published: (2025)
by: Shi, Hao, et al.
Published: (2025)
ReSim: Reliable World Simulation for Autonomous Driving
by: Yang, Jiazhi, et al.
Published: (2025)
by: Yang, Jiazhi, et al.
Published: (2025)
Reconstructive Visual Instruction Tuning
by: Wang, Haochen, et al.
Published: (2024)
by: Wang, Haochen, et al.
Published: (2024)
Multi-GraspLLM: A Multimodal LLM for Multi-Hand Semantic Guided Grasp Generation
by: Li, Haosheng, et al.
Published: (2024)
by: Li, Haosheng, et al.
Published: (2024)
Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback
by: Zhou, Yucheng, et al.
Published: (2025)
by: Zhou, Yucheng, et al.
Published: (2025)
Condition Errors Refinement in Autoregressive Image Generation with Diffusion Loss
by: Zhou, Yucheng, et al.
Published: (2026)
by: Zhou, Yucheng, et al.
Published: (2026)
Sim4Seg: Boosting Multimodal Multi-disease Medical Diagnosis Segmentation with Region-Aware Vision-Language Similarity Masks
by: Song, Lingran, et al.
Published: (2025)
by: Song, Lingran, et al.
Published: (2025)
MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration
by: Zhou, Yucheng, et al.
Published: (2025)
by: Zhou, Yucheng, et al.
Published: (2025)
SegGrasp: Zero-Shot Task-Oriented Grasping via Semantic and Geometric Guided Segmentation
by: Li, Haosheng, et al.
Published: (2024)
by: Li, Haosheng, et al.
Published: (2024)
DrivingSphere: Building a High-fidelity 4D World for Closed-loop Simulation
by: Yan, Tianyi, et al.
Published: (2024)
by: Yan, Tianyi, et al.
Published: (2024)
RLGF: Reinforcement Learning with Geometric Feedback for Autonomous Driving Video Generation
by: Yan, Tianyi, et al.
Published: (2025)
by: Yan, Tianyi, et al.
Published: (2025)
FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving
by: Zeng, Shuang, et al.
Published: (2025)
by: Zeng, Shuang, et al.
Published: (2025)
BFA: Best-Feature-Aware Fusion for Multi-View Fine-grained Manipulation
by: Lan, Zihan, et al.
Published: (2025)
by: Lan, Zihan, et al.
Published: (2025)
Persistent Autoregressive Mapping with Traffic Rules for Autonomous Driving
by: Liang, Shiyi, et al.
Published: (2025)
by: Liang, Shiyi, et al.
Published: (2025)
Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation
by: Zheng, Anlin, et al.
Published: (2025)
by: Zheng, Anlin, et al.
Published: (2025)
GeoVLA: Empowering 3D Representations in Vision-Language-Action Models
by: Sun, Lin, et al.
Published: (2025)
by: Sun, Lin, et al.
Published: (2025)
The impact of baffle and taper channel tilt angle on the output performance of proton‐exchange membrane fuel cells
by: Tiancai Cheng, et al.
Published: (2024)
by: Tiancai Cheng, et al.
Published: (2024)
Editable Scene Simulation for Autonomous Driving via Collaborative LLM-Agents
by: Wei, Yuxi, et al.
Published: (2024)
by: Wei, Yuxi, et al.
Published: (2024)
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
by: Shi, Hao, et al.
Published: (2025)
by: Shi, Hao, et al.
Published: (2025)
RoboMatrix: A Skill-centric Hierarchical Framework for Scalable Robot Task Planning and Execution in Open-World
by: Mao, Weixin, et al.
Published: (2024)
by: Mao, Weixin, et al.
Published: (2024)
TrajDiff: End-to-end Autonomous Driving without Perception Annotation
by: Gui, Xingtai, et al.
Published: (2025)
by: Gui, Xingtai, et al.
Published: (2025)
From Human Intention to Action Prediction: Intention-Driven End-to-End Autonomous Driving
by: Zheng, Huan, et al.
Published: (2025)
by: Zheng, Huan, et al.
Published: (2025)
RoboGSim: A Real2Sim2Real Robotic Gaussian Splatting Simulator
by: Li, Xinhai, et al.
Published: (2024)
by: Li, Xinhai, et al.
Published: (2024)
OLiDM: Object-aware LiDAR Diffusion Models for Autonomous Driving
by: Yan, Tianyi, et al.
Published: (2024)
by: Yan, Tianyi, et al.
Published: (2024)
Continual LLaVA: Continual Instruction Tuning in Large Vision-Language Models
by: Cao, Meng, et al.
Published: (2024)
by: Cao, Meng, et al.
Published: (2024)
Diffusion Model with Representation Alignment for Protein Inverse Folding
by: Wang, Chenglin, et al.
Published: (2024)
by: Wang, Chenglin, et al.
Published: (2024)
Information Coordination as a Bridge: A Neuro-Symbolic Architecture for Reliable Autonomous Driving Scene Understanding
by: Liu, Shuo, et al.
Published: (2026)
by: Liu, Shuo, et al.
Published: (2026)
Merlin:Empowering Multimodal LLMs with Foresight Minds
by: Yu, En, et al.
Published: (2023)
by: Yu, En, et al.
Published: (2023)
Similar Items
-
PADriver: Towards Personalized Autonomous Driving
by: Kou, Genghua, et al.
Published: (2025) -
Language Prompt for Autonomous Driving
by: Wu, Dongming, et al.
Published: (2023) -
SubjectDrive: Scaling Generative Data in Autonomous Driving via Subject Control
by: Huang, Binyuan, et al.
Published: (2024) -
Panacea+: Panoramic and Controllable Video Generation for Autonomous Driving
by: Wen, Yuqing, et al.
Published: (2024) -
Glad: A Streaming Scene Generator for Autonomous Driving
by: Xie, Bin, et al.
Published: (2025)