Is a 3D-Tokenized LLM the Key to Reliable Autonomous Driving?
Fuente:
arXiv
Guardado en:
| Autores principales: | Bai, Yifan, Wu, Dongming, Liu, Yingfei, Jia, Fan, Mao, Weixin, Zhang, Ziheng, Zhao, Yucheng, Shen, Jianbing, Wei, Xing, Wang, Tiancai, Zhang, Xiangyu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PADriver: Towards Personalized Autonomous Driving
por: Kou, Genghua, et al.
Publicado: (2025)
por: Kou, Genghua, et al.
Publicado: (2025)
Language Prompt for Autonomous Driving
por: Wu, Dongming, et al.
Publicado: (2023)
por: Wu, Dongming, et al.
Publicado: (2023)
SubjectDrive: Scaling Generative Data in Autonomous Driving via Subject Control
por: Huang, Binyuan, et al.
Publicado: (2024)
por: Huang, Binyuan, et al.
Publicado: (2024)
Panacea+: Panoramic and Controllable Video Generation for Autonomous Driving
por: Wen, Yuqing, et al.
Publicado: (2024)
por: Wen, Yuqing, et al.
Publicado: (2024)
Glad: A Streaming Scene Generator for Autonomous Driving
por: Xie, Bin, et al.
Publicado: (2025)
por: Xie, Bin, et al.
Publicado: (2025)
Stream Query Denoising for Vectorized HD Map Construction
por: Wang, Shuo, et al.
Publicado: (2024)
por: Wang, Shuo, et al.
Publicado: (2024)
Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness
por: Wang, Haochen, et al.
Publicado: (2025)
por: Wang, Haochen, et al.
Publicado: (2025)
DEGround: An Effective Baseline for Ego-centric 3D Visual Grounding with a Homogeneous Framework
por: Zhang, Yani, et al.
Publicado: (2025)
por: Zhang, Yani, et al.
Publicado: (2025)
Hita: Holistic Tokenizer for Autoregressive Image Generation
por: Zheng, Anlin, et al.
Publicado: (2025)
por: Zheng, Anlin, et al.
Publicado: (2025)
RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping
por: Wu, Dongming, et al.
Publicado: (2025)
por: Wu, Dongming, et al.
Publicado: (2025)
Accelerating Training of Autoregressive Video Generation Models via Local Optimization with Representation Continuity
por: Zhou, Yucheng, et al.
Publicado: (2026)
por: Zhou, Yucheng, et al.
Publicado: (2026)
DME-Driver: Integrating Human Decision Logic and 3D Scene Perception in Autonomous Driving
por: Han, Wencheng, et al.
Publicado: (2024)
por: Han, Wencheng, et al.
Publicado: (2024)
SpatialActor: Exploring Disentangled Spatial Representations for Robust Robotic Manipulation
por: Shi, Hao, et al.
Publicado: (2025)
por: Shi, Hao, et al.
Publicado: (2025)
ReSim: Reliable World Simulation for Autonomous Driving
por: Yang, Jiazhi, et al.
Publicado: (2025)
por: Yang, Jiazhi, et al.
Publicado: (2025)
Reconstructive Visual Instruction Tuning
por: Wang, Haochen, et al.
Publicado: (2024)
por: Wang, Haochen, et al.
Publicado: (2024)
Multi-GraspLLM: A Multimodal LLM for Multi-Hand Semantic Guided Grasp Generation
por: Li, Haosheng, et al.
Publicado: (2024)
por: Li, Haosheng, et al.
Publicado: (2024)
Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback
por: Zhou, Yucheng, et al.
Publicado: (2025)
por: Zhou, Yucheng, et al.
Publicado: (2025)
Condition Errors Refinement in Autoregressive Image Generation with Diffusion Loss
por: Zhou, Yucheng, et al.
Publicado: (2026)
por: Zhou, Yucheng, et al.
Publicado: (2026)
Sim4Seg: Boosting Multimodal Multi-disease Medical Diagnosis Segmentation with Region-Aware Vision-Language Similarity Masks
por: Song, Lingran, et al.
Publicado: (2025)
por: Song, Lingran, et al.
Publicado: (2025)
MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration
por: Zhou, Yucheng, et al.
Publicado: (2025)
por: Zhou, Yucheng, et al.
Publicado: (2025)
SegGrasp: Zero-Shot Task-Oriented Grasping via Semantic and Geometric Guided Segmentation
por: Li, Haosheng, et al.
Publicado: (2024)
por: Li, Haosheng, et al.
Publicado: (2024)
DrivingSphere: Building a High-fidelity 4D World for Closed-loop Simulation
por: Yan, Tianyi, et al.
Publicado: (2024)
por: Yan, Tianyi, et al.
Publicado: (2024)
RLGF: Reinforcement Learning with Geometric Feedback for Autonomous Driving Video Generation
por: Yan, Tianyi, et al.
Publicado: (2025)
por: Yan, Tianyi, et al.
Publicado: (2025)
FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving
por: Zeng, Shuang, et al.
Publicado: (2025)
por: Zeng, Shuang, et al.
Publicado: (2025)
BFA: Best-Feature-Aware Fusion for Multi-View Fine-grained Manipulation
por: Lan, Zihan, et al.
Publicado: (2025)
por: Lan, Zihan, et al.
Publicado: (2025)
Persistent Autoregressive Mapping with Traffic Rules for Autonomous Driving
por: Liang, Shiyi, et al.
Publicado: (2025)
por: Liang, Shiyi, et al.
Publicado: (2025)
Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation
por: Zheng, Anlin, et al.
Publicado: (2025)
por: Zheng, Anlin, et al.
Publicado: (2025)
GeoVLA: Empowering 3D Representations in Vision-Language-Action Models
por: Sun, Lin, et al.
Publicado: (2025)
por: Sun, Lin, et al.
Publicado: (2025)
The impact of baffle and taper channel tilt angle on the output performance of proton‐exchange membrane fuel cells
por: Tiancai Cheng, et al.
Publicado: (2024)
por: Tiancai Cheng, et al.
Publicado: (2024)
Editable Scene Simulation for Autonomous Driving via Collaborative LLM-Agents
por: Wei, Yuxi, et al.
Publicado: (2024)
por: Wei, Yuxi, et al.
Publicado: (2024)
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
por: Shi, Hao, et al.
Publicado: (2025)
por: Shi, Hao, et al.
Publicado: (2025)
RoboMatrix: A Skill-centric Hierarchical Framework for Scalable Robot Task Planning and Execution in Open-World
por: Mao, Weixin, et al.
Publicado: (2024)
por: Mao, Weixin, et al.
Publicado: (2024)
TrajDiff: End-to-end Autonomous Driving without Perception Annotation
por: Gui, Xingtai, et al.
Publicado: (2025)
por: Gui, Xingtai, et al.
Publicado: (2025)
From Human Intention to Action Prediction: Intention-Driven End-to-End Autonomous Driving
por: Zheng, Huan, et al.
Publicado: (2025)
por: Zheng, Huan, et al.
Publicado: (2025)
RoboGSim: A Real2Sim2Real Robotic Gaussian Splatting Simulator
por: Li, Xinhai, et al.
Publicado: (2024)
por: Li, Xinhai, et al.
Publicado: (2024)
OLiDM: Object-aware LiDAR Diffusion Models for Autonomous Driving
por: Yan, Tianyi, et al.
Publicado: (2024)
por: Yan, Tianyi, et al.
Publicado: (2024)
Continual LLaVA: Continual Instruction Tuning in Large Vision-Language Models
por: Cao, Meng, et al.
Publicado: (2024)
por: Cao, Meng, et al.
Publicado: (2024)
Diffusion Model with Representation Alignment for Protein Inverse Folding
por: Wang, Chenglin, et al.
Publicado: (2024)
por: Wang, Chenglin, et al.
Publicado: (2024)
Information Coordination as a Bridge: A Neuro-Symbolic Architecture for Reliable Autonomous Driving Scene Understanding
por: Liu, Shuo, et al.
Publicado: (2026)
por: Liu, Shuo, et al.
Publicado: (2026)
Merlin:Empowering Multimodal LLMs with Foresight Minds
por: Yu, En, et al.
Publicado: (2023)
por: Yu, En, et al.
Publicado: (2023)
Ejemplares similares
-
PADriver: Towards Personalized Autonomous Driving
por: Kou, Genghua, et al.
Publicado: (2025) -
Language Prompt for Autonomous Driving
por: Wu, Dongming, et al.
Publicado: (2023) -
SubjectDrive: Scaling Generative Data in Autonomous Driving via Subject Control
por: Huang, Binyuan, et al.
Publicado: (2024) -
Panacea+: Panoramic and Controllable Video Generation for Autonomous Driving
por: Wen, Yuqing, et al.
Publicado: (2024) -
Glad: A Streaming Scene Generator for Autonomous Driving
por: Xie, Bin, et al.
Publicado: (2025)