Learning from Risk: LLM-Guided Generation of Safety-Critical Scenarios with Prior Knowledge

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yuhang, Huang, Heye, Xu, Zhenhua, Sun, Kailai, Guo, Baoshen, Zhao, Jinhua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909925350834176
author Wang, Yuhang
Huang, Heye
Xu, Zhenhua
Sun, Kailai
Guo, Baoshen
Zhao, Jinhua
author_facet Wang, Yuhang
Huang, Heye
Xu, Zhenhua
Sun, Kailai
Guo, Baoshen
Zhao, Jinhua
contents Autonomous driving faces critical challenges in rare long-tail events and complex multi-agent interactions, which are scarce in real-world data yet essential for robust safety validation. This paper presents a high-fidelity scenario generation framework that integrates a conditional variational autoencoder (CVAE) with a large language model (LLM). The CVAE encodes historical trajectories and map information from large-scale naturalistic datasets to learn latent traffic structures, enabling the generation of physically consistent base scenarios. Building on this, the LLM acts as an adversarial reasoning engine, parsing unstructured scene descriptions into domain-specific loss functions and dynamically guiding scenario generation across varying risk levels. This knowledge-driven optimization balances realism with controllability, ensuring that generated scenarios remain both plausible and risk-sensitive. Extensive experiments in CARLA and SMARTS demonstrate that our framework substantially increases the coverage of high-risk and long-tail events, improves consistency between simulated and real-world traffic distributions, and exposes autonomous driving systems to interactions that are significantly more challenging than those produced by existing rule- or data-driven methods. These results establish a new pathway for safety validation, enabling principled stress-testing of autonomous systems under rare but consequential events.
format Preprint
id arxiv_https___arxiv_org_abs_2511_20726
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning from Risk: LLM-Guided Generation of Safety-Critical Scenarios with Prior Knowledge
Wang, Yuhang
Huang, Heye
Xu, Zhenhua
Sun, Kailai
Guo, Baoshen
Zhao, Jinhua
Machine Learning
Artificial Intelligence
Autonomous driving faces critical challenges in rare long-tail events and complex multi-agent interactions, which are scarce in real-world data yet essential for robust safety validation. This paper presents a high-fidelity scenario generation framework that integrates a conditional variational autoencoder (CVAE) with a large language model (LLM). The CVAE encodes historical trajectories and map information from large-scale naturalistic datasets to learn latent traffic structures, enabling the generation of physically consistent base scenarios. Building on this, the LLM acts as an adversarial reasoning engine, parsing unstructured scene descriptions into domain-specific loss functions and dynamically guiding scenario generation across varying risk levels. This knowledge-driven optimization balances realism with controllability, ensuring that generated scenarios remain both plausible and risk-sensitive. Extensive experiments in CARLA and SMARTS demonstrate that our framework substantially increases the coverage of high-risk and long-tail events, improves consistency between simulated and real-world traffic distributions, and exposes autonomous driving systems to interactions that are significantly more challenging than those produced by existing rule- or data-driven methods. These results establish a new pathway for safety validation, enabling principled stress-testing of autonomous systems under rare but consequential events.
title Learning from Risk: LLM-Guided Generation of Safety-Critical Scenarios with Prior Knowledge
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2511.20726