Enhancing LLM-based Autonomous Driving with Modular Traffic Light and Sign Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Schmidt, Fabian, Nazar, Noushiq Mohammed Kayilan Abdul, Enzweiler, Markus, Valada, Abhinav
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914163033374720
author Schmidt, Fabian
Nazar, Noushiq Mohammed Kayilan Abdul
Enzweiler, Markus
Valada, Abhinav
author_facet Schmidt, Fabian
Nazar, Noushiq Mohammed Kayilan Abdul
Enzweiler, Markus
Valada, Abhinav
contents Large Language Models (LLMs) are increasingly used for decision-making and planning in autonomous driving, showing promising reasoning capabilities and potential to generalize across diverse traffic situations. However, current LLM-based driving agents lack explicit mechanisms to enforce traffic rules and often struggle to reliably detect small, safety-critical objects such as traffic lights and signs. To address this limitation, we introduce TLS-Assist, a modular redundancy layer that augments LLM-based autonomous driving agents with explicit traffic light and sign recognition. TLS-Assist converts detections into structured natural language messages that are injected into the LLM input, enforcing explicit attention to safety-critical cues. The framework is plug-and-play, model-agnostic, and supports both single-view and multi-view camera setups. We evaluate TLS-Assist in a closed-loop setup on the LangAuto benchmark in CARLA. The results demonstrate relative driving performance improvements of up to 14% over LMDrive and 7% over BEVDriver, while consistently reducing traffic light and sign infractions. We publicly release the code and models on https://github.com/iis-esslingen/TLS-Assist.
format Preprint
id arxiv_https___arxiv_org_abs_2511_14391
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing LLM-based Autonomous Driving with Modular Traffic Light and Sign Recognition
Schmidt, Fabian
Nazar, Noushiq Mohammed Kayilan Abdul
Enzweiler, Markus
Valada, Abhinav
Computer Vision and Pattern Recognition
Large Language Models (LLMs) are increasingly used for decision-making and planning in autonomous driving, showing promising reasoning capabilities and potential to generalize across diverse traffic situations. However, current LLM-based driving agents lack explicit mechanisms to enforce traffic rules and often struggle to reliably detect small, safety-critical objects such as traffic lights and signs. To address this limitation, we introduce TLS-Assist, a modular redundancy layer that augments LLM-based autonomous driving agents with explicit traffic light and sign recognition. TLS-Assist converts detections into structured natural language messages that are injected into the LLM input, enforcing explicit attention to safety-critical cues. The framework is plug-and-play, model-agnostic, and supports both single-view and multi-view camera setups. We evaluate TLS-Assist in a closed-loop setup on the LangAuto benchmark in CARLA. The results demonstrate relative driving performance improvements of up to 14% over LMDrive and 7% over BEVDriver, while consistently reducing traffic light and sign infractions. We publicly release the code and models on https://github.com/iis-esslingen/TLS-Assist.
title Enhancing LLM-based Autonomous Driving with Modular Traffic Light and Sign Recognition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.14391