VILTA: A VLM-in-the-Loop Adversary for Enhancing Driving Policy Robustness

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Qimao, Li, Fang, Xu, Shaoqing, Lai, Zhiyi, Xie, Zixun, Luo, Yuechen, Jiang, Shengyin, Li, Hanbing, Chen, Long, Wang, Bing, Zhang, Yi, Yang, Zhi-Xin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914263296114688
author Chen, Qimao
Li, Fang
Xu, Shaoqing
Lai, Zhiyi
Xie, Zixun
Luo, Yuechen
Jiang, Shengyin
Li, Hanbing
Chen, Long
Wang, Bing
Zhang, Yi
Yang, Zhi-Xin
author_facet Chen, Qimao
Li, Fang
Xu, Shaoqing
Lai, Zhiyi
Xie, Zixun
Luo, Yuechen
Jiang, Shengyin
Li, Hanbing
Chen, Long
Wang, Bing
Zhang, Yi
Yang, Zhi-Xin
contents The safe deployment of autonomous driving (AD) systems is fundamentally hindered by the long-tail problem, where rare yet critical driving scenarios are severely underrepresented in real-world data. Existing solutions including safety-critical scenario generation and closed-loop learning often rely on rule-based heuristics, resampling methods and generative models learned from offline datasets, limiting their ability to produce diverse and novel challenges. While recent works leverage Vision Language Models (VLMs) to produce scene descriptions that guide a separate, downstream model in generating hazardous trajectories for agents, such two-stage framework constrains the generative potential of VLMs, as the diversity of the final trajectories is ultimately limited by the generalization ceiling of the downstream algorithm. To overcome these limitations, we introduce VILTA (VLM-In-the-Loop Trajectory Adversary), a novel framework that integrates a VLM into the closed-loop training of AD agents. Unlike prior works, VILTA actively participates in the training loop by comprehending the dynamic driving environment and strategically generating challenging scenarios through direct, fine-grained editing of surrounding agents' future trajectories. This direct-editing approach fully leverages the VLM's powerful generalization capabilities to create a diverse curriculum of plausible yet challenging scenarios that extend beyond the scope of traditional methods. We demonstrate that our approach substantially enhances the safety and robustness of the resulting AD policy, particularly in its ability to navigate critical long-tail events.
format Preprint
id arxiv_https___arxiv_org_abs_2601_12672
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VILTA: A VLM-in-the-Loop Adversary for Enhancing Driving Policy Robustness
Chen, Qimao
Li, Fang
Xu, Shaoqing
Lai, Zhiyi
Xie, Zixun
Luo, Yuechen
Jiang, Shengyin
Li, Hanbing
Chen, Long
Wang, Bing
Zhang, Yi
Yang, Zhi-Xin
Computer Vision and Pattern Recognition
The safe deployment of autonomous driving (AD) systems is fundamentally hindered by the long-tail problem, where rare yet critical driving scenarios are severely underrepresented in real-world data. Existing solutions including safety-critical scenario generation and closed-loop learning often rely on rule-based heuristics, resampling methods and generative models learned from offline datasets, limiting their ability to produce diverse and novel challenges. While recent works leverage Vision Language Models (VLMs) to produce scene descriptions that guide a separate, downstream model in generating hazardous trajectories for agents, such two-stage framework constrains the generative potential of VLMs, as the diversity of the final trajectories is ultimately limited by the generalization ceiling of the downstream algorithm. To overcome these limitations, we introduce VILTA (VLM-In-the-Loop Trajectory Adversary), a novel framework that integrates a VLM into the closed-loop training of AD agents. Unlike prior works, VILTA actively participates in the training loop by comprehending the dynamic driving environment and strategically generating challenging scenarios through direct, fine-grained editing of surrounding agents' future trajectories. This direct-editing approach fully leverages the VLM's powerful generalization capabilities to create a diverse curriculum of plausible yet challenging scenarios that extend beyond the scope of traditional methods. We demonstrate that our approach substantially enhances the safety and robustness of the resulting AD policy, particularly in its ability to navigate critical long-tail events.
title VILTA: A VLM-in-the-Loop Adversary for Enhancing Driving Policy Robustness
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.12672