ElectroVizQA: How well do Multi-modal LLMs perform in Electronics Visual Question Answering?

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Meshram, Pragati Shuddhodhan, Karthikeyan, Swetha, Bhavya, Bhavya, Bhat, Suma
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912573656399872
author Meshram, Pragati Shuddhodhan
Karthikeyan, Swetha
Bhavya, Bhavya
Bhat, Suma
author_facet Meshram, Pragati Shuddhodhan
Karthikeyan, Swetha
Bhavya, Bhavya
Bhat, Suma
contents Multi-modal Large Language Models (MLLMs) are gaining significant attention for their ability to process multi-modal data, providing enhanced contextual understanding of complex problems. MLLMs have demonstrated exceptional capabilities in tasks such as Visual Question Answering (VQA); however, they often struggle with fundamental engineering problems, and there is a scarcity of specialized datasets for training on topics like digital electronics. To address this gap, we propose a benchmark dataset called ElectroVizQA specifically designed to evaluate MLLMs' performance on digital electronic circuit problems commonly found in undergraduate curricula. This dataset, the first of its kind tailored for the VQA task in digital electronics, comprises approximately 626 visual questions, offering a comprehensive overview of digital electronics topics. This paper rigorously assesses the extent to which MLLMs can understand and solve digital electronic circuit questions, providing insights into their capabilities and limitations within this specialized domain. By introducing this benchmark dataset, we aim to motivate further research and development in the application of MLLMs to engineering education, ultimately bridging the performance gap and enhancing the efficacy of these models in technical fields.
format Preprint
id arxiv_https___arxiv_org_abs_2412_00102
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ElectroVizQA: How well do Multi-modal LLMs perform in Electronics Visual Question Answering?
Meshram, Pragati Shuddhodhan
Karthikeyan, Swetha
Bhavya, Bhavya
Bhat, Suma
Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
Multi-modal Large Language Models (MLLMs) are gaining significant attention for their ability to process multi-modal data, providing enhanced contextual understanding of complex problems. MLLMs have demonstrated exceptional capabilities in tasks such as Visual Question Answering (VQA); however, they often struggle with fundamental engineering problems, and there is a scarcity of specialized datasets for training on topics like digital electronics. To address this gap, we propose a benchmark dataset called ElectroVizQA specifically designed to evaluate MLLMs' performance on digital electronic circuit problems commonly found in undergraduate curricula. This dataset, the first of its kind tailored for the VQA task in digital electronics, comprises approximately 626 visual questions, offering a comprehensive overview of digital electronics topics. This paper rigorously assesses the extent to which MLLMs can understand and solve digital electronic circuit questions, providing insights into their capabilities and limitations within this specialized domain. By introducing this benchmark dataset, we aim to motivate further research and development in the application of MLLMs to engineering education, ultimately bridging the performance gap and enhancing the efficacy of these models in technical fields.
title ElectroVizQA: How well do Multi-modal LLMs perform in Electronics Visual Question Answering?
topic Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
url https://arxiv.org/abs/2412.00102