FLARE: Learning Future-Aware Latent Representations from Vision-Language Models for Autonomous Driving

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xie, Chengen, Sima, Chonghao, Li, Tianyu, Sun, Bin, Wu, Junjie, Hao, Zhihui, Li, Hongyang
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911497818472448
author Xie, Chengen
Sima, Chonghao
Li, Tianyu
Sun, Bin
Wu, Junjie
Hao, Zhihui
Li, Hongyang
author_facet Xie, Chengen
Sima, Chonghao
Li, Tianyu
Sun, Bin
Wu, Junjie
Hao, Zhihui
Li, Hongyang
contents While Vision-Language Models (VLMs) offer rich world knowledge for end-to-end autonomous driving, current approaches heavily rely on labor-intensive language annotations (e.g., VQA) to bridge perception and control. This paradigm suffers from a fundamental mismatch between discrete linguistic tokens and continuous driving trajectories, often leading to suboptimal control policies and inefficient utilization of pre-trained knowledge. To address these challenges, we propose FLARE (Future-aware LAtent REpresentation), a novel framework that activates the visual-semantic capabilities of pre-trained VLMs without requiring language supervision. Instead of aligning with text, we introduce a self-supervised future feature prediction objective. This mechanism compels the model to anticipate scene dynamics and ego-motion directly in the latent space, enabling the learning of robust driving representations from large-scale unlabeled trajectory data. Furthermore, we integrate Group Relative Policy Optimization (GRPO) into the planning process to refine decision-making quality. Extensive experiments on the NAVSIM benchmark demonstrate that FLARE achieves state-of-the-art performance, validating the effectiveness of leveraging VLM knowledge via predictive self-supervision rather than explicit language generation.
format Preprint
id arxiv_https___arxiv_org_abs_2601_05611
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FLARE: Learning Future-Aware Latent Representations from Vision-Language Models for Autonomous Driving
Xie, Chengen
Sima, Chonghao
Li, Tianyu
Sun, Bin
Wu, Junjie
Hao, Zhihui
Li, Hongyang
Computer Vision and Pattern Recognition
While Vision-Language Models (VLMs) offer rich world knowledge for end-to-end autonomous driving, current approaches heavily rely on labor-intensive language annotations (e.g., VQA) to bridge perception and control. This paradigm suffers from a fundamental mismatch between discrete linguistic tokens and continuous driving trajectories, often leading to suboptimal control policies and inefficient utilization of pre-trained knowledge. To address these challenges, we propose FLARE (Future-aware LAtent REpresentation), a novel framework that activates the visual-semantic capabilities of pre-trained VLMs without requiring language supervision. Instead of aligning with text, we introduce a self-supervised future feature prediction objective. This mechanism compels the model to anticipate scene dynamics and ego-motion directly in the latent space, enabling the learning of robust driving representations from large-scale unlabeled trajectory data. Furthermore, we integrate Group Relative Policy Optimization (GRPO) into the planning process to refine decision-making quality. Extensive experiments on the NAVSIM benchmark demonstrate that FLARE achieves state-of-the-art performance, validating the effectiveness of leveraging VLM knowledge via predictive self-supervision rather than explicit language generation.
title FLARE: Learning Future-Aware Latent Representations from Vision-Language Models for Autonomous Driving
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.05611