SignRAG: A Retrieval-Augmented System for Scalable Zero-Shot Road Sign Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Minghao, Zhang, Zhihao, Sidhu, Anmol, Redmill, Keith
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914201645088768
author Zhu, Minghao
Zhang, Zhihao
Sidhu, Anmol
Redmill, Keith
author_facet Zhu, Minghao
Zhang, Zhihao
Sidhu, Anmol
Redmill, Keith
contents Automated road sign recognition is a critical task for intelligent transportation systems, but traditional deep learning methods struggle with the sheer number of sign classes and the impracticality of creating exhaustive labeled datasets. This paper introduces a novel zero-shot recognition framework that adapts the Retrieval-Augmented Generation (RAG) paradigm to address this challenge. Our method first uses a Vision Language Model (VLM) to generate a textual description of a sign from an input image. This description is used to retrieve a small set of the most relevant sign candidates from a vector database of reference designs. Subsequently, a Large Language Model (LLM) reasons over the retrieved candidates to make a final, fine-grained recognition. We validate this approach on a comprehensive set of 303 regulatory signs from the Ohio MUTCD. Experimental results demonstrate the framework's effectiveness, achieving 95.58% accuracy on ideal reference images and 82.45% on challenging real-world road data. This work demonstrates the viability of RAG-based architectures for creating scalable and accurate systems for road sign recognition without task-specific training.
format Preprint
id arxiv_https___arxiv_org_abs_2512_12885
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SignRAG: A Retrieval-Augmented System for Scalable Zero-Shot Road Sign Recognition
Zhu, Minghao
Zhang, Zhihao
Sidhu, Anmol
Redmill, Keith
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Information Retrieval
Robotics
Automated road sign recognition is a critical task for intelligent transportation systems, but traditional deep learning methods struggle with the sheer number of sign classes and the impracticality of creating exhaustive labeled datasets. This paper introduces a novel zero-shot recognition framework that adapts the Retrieval-Augmented Generation (RAG) paradigm to address this challenge. Our method first uses a Vision Language Model (VLM) to generate a textual description of a sign from an input image. This description is used to retrieve a small set of the most relevant sign candidates from a vector database of reference designs. Subsequently, a Large Language Model (LLM) reasons over the retrieved candidates to make a final, fine-grained recognition. We validate this approach on a comprehensive set of 303 regulatory signs from the Ohio MUTCD. Experimental results demonstrate the framework's effectiveness, achieving 95.58% accuracy on ideal reference images and 82.45% on challenging real-world road data. This work demonstrates the viability of RAG-based architectures for creating scalable and accurate systems for road sign recognition without task-specific training.
title SignRAG: A Retrieval-Augmented System for Scalable Zero-Shot Road Sign Recognition
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Information Retrieval
Robotics
url https://arxiv.org/abs/2512.12885