Saved in:
Bibliographic Details
Main Author: Ghosh, Sanjukta
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2510.03376
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915532624625664
author Ghosh, Sanjukta
author_facet Ghosh, Sanjukta
contents Industrial diagrams such as piping and instrumentation diagrams (P&IDs) are essential for the design, operation, and maintenance of industrial plants. Converting these diagrams into digital form is an important step toward building digital twins and enabling intelligent industrial automation. A central challenge in this digitalization process is accurate object detection. Although recent advances have significantly improved object detection algorithms, there remains a lack of methods to automatically evaluate the quality of their outputs. This paper addresses this gap by introducing a framework that employs Visual Language Models (VLMs) to assess object detection results and guide their refinement. The approach exploits the multimodal capabilities of VLMs to identify missing or inconsistent detections, thereby enabling automated quality assessment and improving overall detection performance on complex industrial diagrams.
format Preprint
id arxiv_https___arxiv_org_abs_2510_03376
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Visual Language Model as a Judge for Object Detection in Industrial Diagrams
Ghosh, Sanjukta
Computer Vision and Pattern Recognition
Image and Video Processing
Industrial diagrams such as piping and instrumentation diagrams (P&IDs) are essential for the design, operation, and maintenance of industrial plants. Converting these diagrams into digital form is an important step toward building digital twins and enabling intelligent industrial automation. A central challenge in this digitalization process is accurate object detection. Although recent advances have significantly improved object detection algorithms, there remains a lack of methods to automatically evaluate the quality of their outputs. This paper addresses this gap by introducing a framework that employs Visual Language Models (VLMs) to assess object detection results and guide their refinement. The approach exploits the multimodal capabilities of VLMs to identify missing or inconsistent detections, thereby enabling automated quality assessment and improving overall detection performance on complex industrial diagrams.
title Visual Language Model as a Judge for Object Detection in Industrial Diagrams
topic Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2510.03376