InsightDrive: Insight Scene Representation for End-to-End Autonomous Driving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Ruiqi, Guo, Xianda, Peng, Yanlun, Wei, Qinggong, Wu, Hangbin, Chen, Long
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909935624781824
author Song, Ruiqi
Guo, Xianda
Peng, Yanlun
Wei, Qinggong
Wu, Hangbin
Chen, Long
author_facet Song, Ruiqi
Guo, Xianda
Peng, Yanlun
Wei, Qinggong
Wu, Hangbin
Chen, Long
contents Conventional end-to-end autonomous driving methods often rely on explicit global scene representations, which typically consist of 3D object detection, online mapping, and motion prediction. In contrast, human drivers selectively attend to task-relevant regions and implicitly reason over the broader traffic context. Motivated by this observation, we introduce a lightweight end-to-end autonomous driving framework, InsightDrive. Unlike approaches that directly embed large language models (LLMs), InsightDrive introduces an Insight scene representation that jointly models attention-centric explicit scene representation and reasoning-centric implicit scene representation, so that scene understanding aligns more closely with human cognitive patterns for trajectory planning. To this end, we employ Chain-of-Thought (CoT) instructions to model human driving cognition and design a task-level Mixture-of-Experts (MoE) adapter that injects this knowledge into the autonomous driving model at negligible parameter cost. We further condition the planner on both explicit and implicit scene representations and employ a diffusion-based generative policy, which produces robust trajectory predictions and decisions. The overall framework establishes a knowledge distillation pipeline that transfers human driving knowledge to LLMs and subsequently to onboard models. Extensive experiments on the nuScenes and Navsim benchmarks demonstrate that InsightDrive achieves significant improvements over conventional scene representation approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2503_13047
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle InsightDrive: Insight Scene Representation for End-to-End Autonomous Driving
Song, Ruiqi
Guo, Xianda
Peng, Yanlun
Wei, Qinggong
Wu, Hangbin
Chen, Long
Computer Vision and Pattern Recognition
Conventional end-to-end autonomous driving methods often rely on explicit global scene representations, which typically consist of 3D object detection, online mapping, and motion prediction. In contrast, human drivers selectively attend to task-relevant regions and implicitly reason over the broader traffic context. Motivated by this observation, we introduce a lightweight end-to-end autonomous driving framework, InsightDrive. Unlike approaches that directly embed large language models (LLMs), InsightDrive introduces an Insight scene representation that jointly models attention-centric explicit scene representation and reasoning-centric implicit scene representation, so that scene understanding aligns more closely with human cognitive patterns for trajectory planning. To this end, we employ Chain-of-Thought (CoT) instructions to model human driving cognition and design a task-level Mixture-of-Experts (MoE) adapter that injects this knowledge into the autonomous driving model at negligible parameter cost. We further condition the planner on both explicit and implicit scene representations and employ a diffusion-based generative policy, which produces robust trajectory predictions and decisions. The overall framework establishes a knowledge distillation pipeline that transfers human driving knowledge to LLMs and subsequently to onboard models. Extensive experiments on the nuScenes and Navsim benchmarks demonstrate that InsightDrive achieves significant improvements over conventional scene representation approaches.
title InsightDrive: Insight Scene Representation for End-to-End Autonomous Driving
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.13047