Saved in:
Bibliographic Details
Main Authors: Gu, Chao, Lin, Ke, Luo, Yiyang, Hou, Jiahui, Li, Xiang-Yang
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2409.00909
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910586548256768
author Gu, Chao
Lin, Ke
Luo, Yiyang
Hou, Jiahui
Li, Xiang-Yang
author_facet Gu, Chao
Lin, Ke
Luo, Yiyang
Hou, Jiahui
Li, Xiang-Yang
contents To accurately understand engineering drawings, it is essential to establish the correspondence between images and their description tables within the drawings. Existing document understanding methods predominantly focus on text as the main modality, which is not suitable for documents containing substantial image information. In the field of visual relation detection, the structure of the task inherently limits its capacity to assess relationships among all entity pairs in the drawings. To address this issue, we propose a vision-based relation detection model, named ViRED, to identify the associations between tables and circuits in electrical engineering drawings. Our model mainly consists of three parts: a vision encoder, an object encoder, and a relation decoder. We implement ViRED using PyTorch to evaluate its performance. To validate the efficacy of ViRED, we conduct a series of experiments. The experimental results indicate that, within the engineering drawing dataset, our approach attained an accuracy of 96\% in the task of relation prediction, marking a substantial improvement over existing methodologies. The results also show that ViRED can inference at a fast speed even when there are numerous objects in a single engineering drawing.
format Preprint
id arxiv_https___arxiv_org_abs_2409_00909
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ViRED: Prediction of Visual Relations in Engineering Drawings
Gu, Chao
Lin, Ke
Luo, Yiyang
Hou, Jiahui
Li, Xiang-Yang
Computer Vision and Pattern Recognition
Artificial Intelligence
To accurately understand engineering drawings, it is essential to establish the correspondence between images and their description tables within the drawings. Existing document understanding methods predominantly focus on text as the main modality, which is not suitable for documents containing substantial image information. In the field of visual relation detection, the structure of the task inherently limits its capacity to assess relationships among all entity pairs in the drawings. To address this issue, we propose a vision-based relation detection model, named ViRED, to identify the associations between tables and circuits in electrical engineering drawings. Our model mainly consists of three parts: a vision encoder, an object encoder, and a relation decoder. We implement ViRED using PyTorch to evaluate its performance. To validate the efficacy of ViRED, we conduct a series of experiments. The experimental results indicate that, within the engineering drawing dataset, our approach attained an accuracy of 96\% in the task of relation prediction, marking a substantial improvement over existing methodologies. The results also show that ViRED can inference at a fast speed even when there are numerous objects in a single engineering drawing.
title ViRED: Prediction of Visual Relations in Engineering Drawings
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2409.00909