A Training-Free Framework for Video License Plate Tracking and Recognition with Only One-Shot

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ding, Haoxuan, Wang, Qi, Gao, Junyu, Li, Qiang
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929455657648128
author Ding, Haoxuan
Wang, Qi
Gao, Junyu
Li, Qiang
author_facet Ding, Haoxuan
Wang, Qi
Gao, Junyu
Li, Qiang
contents Traditional license plate detection and recognition models are often trained on closed datasets, limiting their ability to handle the diverse license plate formats across different regions. The emergence of large-scale pre-trained models has shown exceptional generalization capabilities, enabling few-shot and zero-shot learning. We propose OneShotLP, a training-free framework for video-based license plate detection and recognition, leveraging these advanced models. Starting with the license plate position in the first video frame, our method tracks this position across subsequent frames using a point tracking module, creating a trajectory of prompts. These prompts are input into a segmentation module that uses a promptable large segmentation model to generate local masks of the license plate regions. The segmented areas are then processed by multimodal large language models (MLLMs) for accurate license plate recognition. OneShotLP offers significant advantages, including the ability to function effectively without extensive training data and adaptability to various license plate styles. Experimental results on UFPR-ALPR and SSIG-SegPlate datasets demonstrate the superior accuracy of our approach compared to traditional methods. This highlights the potential of leveraging pre-trained models for diverse real-world applications in intelligent transportation systems. The code is available at https://github.com/Dinghaoxuan/OneShotLP.
format Preprint
id arxiv_https___arxiv_org_abs_2408_05729
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Training-Free Framework for Video License Plate Tracking and Recognition with Only One-Shot
Ding, Haoxuan
Wang, Qi
Gao, Junyu
Li, Qiang
Computer Vision and Pattern Recognition
Traditional license plate detection and recognition models are often trained on closed datasets, limiting their ability to handle the diverse license plate formats across different regions. The emergence of large-scale pre-trained models has shown exceptional generalization capabilities, enabling few-shot and zero-shot learning. We propose OneShotLP, a training-free framework for video-based license plate detection and recognition, leveraging these advanced models. Starting with the license plate position in the first video frame, our method tracks this position across subsequent frames using a point tracking module, creating a trajectory of prompts. These prompts are input into a segmentation module that uses a promptable large segmentation model to generate local masks of the license plate regions. The segmented areas are then processed by multimodal large language models (MLLMs) for accurate license plate recognition. OneShotLP offers significant advantages, including the ability to function effectively without extensive training data and adaptability to various license plate styles. Experimental results on UFPR-ALPR and SSIG-SegPlate datasets demonstrate the superior accuracy of our approach compared to traditional methods. This highlights the potential of leveraging pre-trained models for diverse real-world applications in intelligent transportation systems. The code is available at https://github.com/Dinghaoxuan/OneShotLP.
title A Training-Free Framework for Video License Plate Tracking and Recognition with Only One-Shot
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2408.05729