Is a 3D-Tokenized LLM the Key to Reliable Autonomous Driving?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bai, Yifan, Wu, Dongming, Liu, Yingfei, Jia, Fan, Mao, Weixin, Zhang, Ziheng, Zhao, Yucheng, Shen, Jianbing, Wei, Xing, Wang, Tiancai, Zhang, Xiangyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PADriver: Towards Personalized Autonomous Driving
von: Kou, Genghua, et al.
Veröffentlicht: (2025)
von: Kou, Genghua, et al.
Veröffentlicht: (2025)
Language Prompt for Autonomous Driving
von: Wu, Dongming, et al.
Veröffentlicht: (2023)
von: Wu, Dongming, et al.
Veröffentlicht: (2023)
SubjectDrive: Scaling Generative Data in Autonomous Driving via Subject Control
von: Huang, Binyuan, et al.
Veröffentlicht: (2024)
von: Huang, Binyuan, et al.
Veröffentlicht: (2024)
Panacea+: Panoramic and Controllable Video Generation for Autonomous Driving
von: Wen, Yuqing, et al.
Veröffentlicht: (2024)
von: Wen, Yuqing, et al.
Veröffentlicht: (2024)
Glad: A Streaming Scene Generator for Autonomous Driving
von: Xie, Bin, et al.
Veröffentlicht: (2025)
von: Xie, Bin, et al.
Veröffentlicht: (2025)
Stream Query Denoising for Vectorized HD Map Construction
von: Wang, Shuo, et al.
Veröffentlicht: (2024)
von: Wang, Shuo, et al.
Veröffentlicht: (2024)
Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
DEGround: An Effective Baseline for Ego-centric 3D Visual Grounding with a Homogeneous Framework
von: Zhang, Yani, et al.
Veröffentlicht: (2025)
von: Zhang, Yani, et al.
Veröffentlicht: (2025)
Hita: Holistic Tokenizer for Autoregressive Image Generation
von: Zheng, Anlin, et al.
Veröffentlicht: (2025)
von: Zheng, Anlin, et al.
Veröffentlicht: (2025)
RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping
von: Wu, Dongming, et al.
Veröffentlicht: (2025)
von: Wu, Dongming, et al.
Veröffentlicht: (2025)
Accelerating Training of Autoregressive Video Generation Models via Local Optimization with Representation Continuity
von: Zhou, Yucheng, et al.
Veröffentlicht: (2026)
von: Zhou, Yucheng, et al.
Veröffentlicht: (2026)
DME-Driver: Integrating Human Decision Logic and 3D Scene Perception in Autonomous Driving
von: Han, Wencheng, et al.
Veröffentlicht: (2024)
von: Han, Wencheng, et al.
Veröffentlicht: (2024)
SpatialActor: Exploring Disentangled Spatial Representations for Robust Robotic Manipulation
von: Shi, Hao, et al.
Veröffentlicht: (2025)
von: Shi, Hao, et al.
Veröffentlicht: (2025)
ReSim: Reliable World Simulation for Autonomous Driving
von: Yang, Jiazhi, et al.
Veröffentlicht: (2025)
von: Yang, Jiazhi, et al.
Veröffentlicht: (2025)
Reconstructive Visual Instruction Tuning
von: Wang, Haochen, et al.
Veröffentlicht: (2024)
von: Wang, Haochen, et al.
Veröffentlicht: (2024)
Multi-GraspLLM: A Multimodal LLM for Multi-Hand Semantic Guided Grasp Generation
von: Li, Haosheng, et al.
Veröffentlicht: (2024)
von: Li, Haosheng, et al.
Veröffentlicht: (2024)
Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback
von: Zhou, Yucheng, et al.
Veröffentlicht: (2025)
von: Zhou, Yucheng, et al.
Veröffentlicht: (2025)
Condition Errors Refinement in Autoregressive Image Generation with Diffusion Loss
von: Zhou, Yucheng, et al.
Veröffentlicht: (2026)
von: Zhou, Yucheng, et al.
Veröffentlicht: (2026)
Sim4Seg: Boosting Multimodal Multi-disease Medical Diagnosis Segmentation with Region-Aware Vision-Language Similarity Masks
von: Song, Lingran, et al.
Veröffentlicht: (2025)
von: Song, Lingran, et al.
Veröffentlicht: (2025)
MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration
von: Zhou, Yucheng, et al.
Veröffentlicht: (2025)
von: Zhou, Yucheng, et al.
Veröffentlicht: (2025)
SegGrasp: Zero-Shot Task-Oriented Grasping via Semantic and Geometric Guided Segmentation
von: Li, Haosheng, et al.
Veröffentlicht: (2024)
von: Li, Haosheng, et al.
Veröffentlicht: (2024)
DrivingSphere: Building a High-fidelity 4D World for Closed-loop Simulation
von: Yan, Tianyi, et al.
Veröffentlicht: (2024)
von: Yan, Tianyi, et al.
Veröffentlicht: (2024)
RLGF: Reinforcement Learning with Geometric Feedback for Autonomous Driving Video Generation
von: Yan, Tianyi, et al.
Veröffentlicht: (2025)
von: Yan, Tianyi, et al.
Veröffentlicht: (2025)
FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving
von: Zeng, Shuang, et al.
Veröffentlicht: (2025)
von: Zeng, Shuang, et al.
Veröffentlicht: (2025)
BFA: Best-Feature-Aware Fusion for Multi-View Fine-grained Manipulation
von: Lan, Zihan, et al.
Veröffentlicht: (2025)
von: Lan, Zihan, et al.
Veröffentlicht: (2025)
Persistent Autoregressive Mapping with Traffic Rules for Autonomous Driving
von: Liang, Shiyi, et al.
Veröffentlicht: (2025)
von: Liang, Shiyi, et al.
Veröffentlicht: (2025)
Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation
von: Zheng, Anlin, et al.
Veröffentlicht: (2025)
von: Zheng, Anlin, et al.
Veröffentlicht: (2025)
GeoVLA: Empowering 3D Representations in Vision-Language-Action Models
von: Sun, Lin, et al.
Veröffentlicht: (2025)
von: Sun, Lin, et al.
Veröffentlicht: (2025)
The impact of baffle and taper channel tilt angle on the output performance of proton‐exchange membrane fuel cells
von: Tiancai Cheng, et al.
Veröffentlicht: (2024)
von: Tiancai Cheng, et al.
Veröffentlicht: (2024)
Editable Scene Simulation for Autonomous Driving via Collaborative LLM-Agents
von: Wei, Yuxi, et al.
Veröffentlicht: (2024)
von: Wei, Yuxi, et al.
Veröffentlicht: (2024)
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
von: Shi, Hao, et al.
Veröffentlicht: (2025)
von: Shi, Hao, et al.
Veröffentlicht: (2025)
RoboMatrix: A Skill-centric Hierarchical Framework for Scalable Robot Task Planning and Execution in Open-World
von: Mao, Weixin, et al.
Veröffentlicht: (2024)
von: Mao, Weixin, et al.
Veröffentlicht: (2024)
TrajDiff: End-to-end Autonomous Driving without Perception Annotation
von: Gui, Xingtai, et al.
Veröffentlicht: (2025)
von: Gui, Xingtai, et al.
Veröffentlicht: (2025)
From Human Intention to Action Prediction: Intention-Driven End-to-End Autonomous Driving
von: Zheng, Huan, et al.
Veröffentlicht: (2025)
von: Zheng, Huan, et al.
Veröffentlicht: (2025)
RoboGSim: A Real2Sim2Real Robotic Gaussian Splatting Simulator
von: Li, Xinhai, et al.
Veröffentlicht: (2024)
von: Li, Xinhai, et al.
Veröffentlicht: (2024)
OLiDM: Object-aware LiDAR Diffusion Models for Autonomous Driving
von: Yan, Tianyi, et al.
Veröffentlicht: (2024)
von: Yan, Tianyi, et al.
Veröffentlicht: (2024)
Continual LLaVA: Continual Instruction Tuning in Large Vision-Language Models
von: Cao, Meng, et al.
Veröffentlicht: (2024)
von: Cao, Meng, et al.
Veröffentlicht: (2024)
Diffusion Model with Representation Alignment for Protein Inverse Folding
von: Wang, Chenglin, et al.
Veröffentlicht: (2024)
von: Wang, Chenglin, et al.
Veröffentlicht: (2024)
Information Coordination as a Bridge: A Neuro-Symbolic Architecture for Reliable Autonomous Driving Scene Understanding
von: Liu, Shuo, et al.
Veröffentlicht: (2026)
von: Liu, Shuo, et al.
Veröffentlicht: (2026)
Merlin:Empowering Multimodal LLMs with Foresight Minds
von: Yu, En, et al.
Veröffentlicht: (2023)
von: Yu, En, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
PADriver: Towards Personalized Autonomous Driving
von: Kou, Genghua, et al.
Veröffentlicht: (2025) -
Language Prompt for Autonomous Driving
von: Wu, Dongming, et al.
Veröffentlicht: (2023) -
SubjectDrive: Scaling Generative Data in Autonomous Driving via Subject Control
von: Huang, Binyuan, et al.
Veröffentlicht: (2024) -
Panacea+: Panoramic and Controllable Video Generation for Autonomous Driving
von: Wen, Yuqing, et al.
Veröffentlicht: (2024) -
Glad: A Streaming Scene Generator for Autonomous Driving
von: Xie, Bin, et al.
Veröffentlicht: (2025)