EvoDriveVLA: Evolving Driving VLA Models via Collaborative Perception-Planning Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cao, Jiajun, Zhang, Xiaoan, Wei, Xiaobao, Huang, Liyuqiu, Wang, Zijian, Zhang, Hanzhen, Jia, Zhengyu, Mao, Wei, Wang, Hao, Liu, Xianming, Zhou, Shuchang, Wang, Yang, Zhang, Shanghang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910206842109952
author Cao, Jiajun
Zhang, Xiaoan
Wei, Xiaobao
Huang, Liyuqiu
Wang, Zijian
Zhang, Hanzhen
Jia, Zhengyu
Mao, Wei
Wang, Hao
Liu, Xianming
Zhou, Shuchang
Wang, Yang
Zhang, Shanghang
author_facet Cao, Jiajun
Zhang, Xiaoan
Wei, Xiaobao
Huang, Liyuqiu
Wang, Zijian
Zhang, Hanzhen
Jia, Zhengyu
Mao, Wei
Wang, Hao
Liu, Xianming
Zhou, Shuchang
Wang, Yang
Zhang, Shanghang
contents Vision-Language-Action models have shown great promise for autonomous driving, yet they suffer from degraded perception after unfreezing the visual encoder and struggle with accumulated instability in long-term planning. To address these challenges, we propose EvoDriveVLA-a novel collaborative perception-planning distillation framework that integrates self-anchored perceptual constraints and future-informed trajectory optimization. Specifically, self-anchored visual distillation leverages self-anchor teacher to deliver visual anchoring constraints, regularizing student representations via trajectory-guided key-region awareness. In parallel, future-informed trajectory distillation employs a future-aware oracle teacher with coarse-to-fine trajectory refinement and Monte Carlo dropout sampling to synthesize reasoning trajectories that model future evolutions, enabling the student model to internalize the future-aware insights of the teacher. EvoDriveVLA achieves SOTA performance in nuScenes open-loop evaluation and significantly enhances performance in NAVSIM closed-loop evaluation. Our code is available at: https://github.com/hey-cjj/EvoDriveVLA.
format Preprint
id arxiv_https___arxiv_org_abs_2603_09465
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EvoDriveVLA: Evolving Driving VLA Models via Collaborative Perception-Planning Distillation
Cao, Jiajun
Zhang, Xiaoan
Wei, Xiaobao
Huang, Liyuqiu
Wang, Zijian
Zhang, Hanzhen
Jia, Zhengyu
Mao, Wei
Wang, Hao
Liu, Xianming
Zhou, Shuchang
Wang, Yang
Zhang, Shanghang
Computer Vision and Pattern Recognition
Artificial Intelligence
Vision-Language-Action models have shown great promise for autonomous driving, yet they suffer from degraded perception after unfreezing the visual encoder and struggle with accumulated instability in long-term planning. To address these challenges, we propose EvoDriveVLA-a novel collaborative perception-planning distillation framework that integrates self-anchored perceptual constraints and future-informed trajectory optimization. Specifically, self-anchored visual distillation leverages self-anchor teacher to deliver visual anchoring constraints, regularizing student representations via trajectory-guided key-region awareness. In parallel, future-informed trajectory distillation employs a future-aware oracle teacher with coarse-to-fine trajectory refinement and Monte Carlo dropout sampling to synthesize reasoning trajectories that model future evolutions, enabling the student model to internalize the future-aware insights of the teacher. EvoDriveVLA achieves SOTA performance in nuScenes open-loop evaluation and significantly enhances performance in NAVSIM closed-loop evaluation. Our code is available at: https://github.com/hey-cjj/EvoDriveVLA.
title EvoDriveVLA: Evolving Driving VLA Models via Collaborative Perception-Planning Distillation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2603.09465