EvoDriveVLA: Evolving Driving VLA Models via Collaborative Perception-Planning Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910206842109952 |
|---|---|
| author | Cao, Jiajun Zhang, Xiaoan Wei, Xiaobao Huang, Liyuqiu Wang, Zijian Zhang, Hanzhen Jia, Zhengyu Mao, Wei Wang, Hao Liu, Xianming Zhou, Shuchang Wang, Yang Zhang, Shanghang |
| author_facet | Cao, Jiajun Zhang, Xiaoan Wei, Xiaobao Huang, Liyuqiu Wang, Zijian Zhang, Hanzhen Jia, Zhengyu Mao, Wei Wang, Hao Liu, Xianming Zhou, Shuchang Wang, Yang Zhang, Shanghang |
| contents | Vision-Language-Action models have shown great promise for autonomous driving, yet they suffer from degraded perception after unfreezing the visual encoder and struggle with accumulated instability in long-term planning. To address these challenges, we propose EvoDriveVLA-a novel collaborative perception-planning distillation framework that integrates self-anchored perceptual constraints and future-informed trajectory optimization. Specifically, self-anchored visual distillation leverages self-anchor teacher to deliver visual anchoring constraints, regularizing student representations via trajectory-guided key-region awareness. In parallel, future-informed trajectory distillation employs a future-aware oracle teacher with coarse-to-fine trajectory refinement and Monte Carlo dropout sampling to synthesize reasoning trajectories that model future evolutions, enabling the student model to internalize the future-aware insights of the teacher. EvoDriveVLA achieves SOTA performance in nuScenes open-loop evaluation and significantly enhances performance in NAVSIM closed-loop evaluation. Our code is available at: https://github.com/hey-cjj/EvoDriveVLA. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_09465 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | EvoDriveVLA: Evolving Driving VLA Models via Collaborative Perception-Planning Distillation Cao, Jiajun Zhang, Xiaoan Wei, Xiaobao Huang, Liyuqiu Wang, Zijian Zhang, Hanzhen Jia, Zhengyu Mao, Wei Wang, Hao Liu, Xianming Zhou, Shuchang Wang, Yang Zhang, Shanghang Computer Vision and Pattern Recognition Artificial Intelligence Vision-Language-Action models have shown great promise for autonomous driving, yet they suffer from degraded perception after unfreezing the visual encoder and struggle with accumulated instability in long-term planning. To address these challenges, we propose EvoDriveVLA-a novel collaborative perception-planning distillation framework that integrates self-anchored perceptual constraints and future-informed trajectory optimization. Specifically, self-anchored visual distillation leverages self-anchor teacher to deliver visual anchoring constraints, regularizing student representations via trajectory-guided key-region awareness. In parallel, future-informed trajectory distillation employs a future-aware oracle teacher with coarse-to-fine trajectory refinement and Monte Carlo dropout sampling to synthesize reasoning trajectories that model future evolutions, enabling the student model to internalize the future-aware insights of the teacher. EvoDriveVLA achieves SOTA performance in nuScenes open-loop evaluation and significantly enhances performance in NAVSIM closed-loop evaluation. Our code is available at: https://github.com/hey-cjj/EvoDriveVLA. |
| title | EvoDriveVLA: Evolving Driving VLA Models via Collaborative Perception-Planning Distillation |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2603.09465 |