Solution for SMART-101 Challenge of CVPR Multi-modal Algorithmic Reasoning Task 2024

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ahn, Jinwoo, Park, Junhyeok, Kim, Min-Jun, Kim, Kang-Hyeon, Sohn, So-Yeong, Lee, Yun-Ji, Chang, Du-Seong, Heo, Yu-Jung, Kim, Eun-Sol
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910479297806336
author Ahn, Jinwoo
Park, Junhyeok
Kim, Min-Jun
Kim, Kang-Hyeon
Sohn, So-Yeong
Lee, Yun-Ji
Chang, Du-Seong
Heo, Yu-Jung
Kim, Eun-Sol
author_facet Ahn, Jinwoo
Park, Junhyeok
Kim, Min-Jun
Kim, Kang-Hyeon
Sohn, So-Yeong
Lee, Yun-Ji
Chang, Du-Seong
Heo, Yu-Jung
Kim, Eun-Sol
contents In this paper, the solution of HYU MLLAB KT Team to the Multimodal Algorithmic Reasoning Task: SMART-101 CVPR 2024 Challenge is presented. Beyond conventional visual question-answering problems, the SMART-101 challenge aims to achieve human-level multimodal understanding by tackling complex visio-linguistic puzzles designed for children in the 6-8 age group. To solve this problem, we suggest two main ideas. First, to utilize the reasoning ability of a large-scale language model (LLM), the given visual cues (images) are grounded in the text modality. For this purpose, we generate highly detailed text captions that describe the context of the image and use these captions as input for the LLM. Second, due to the nature of puzzle images, which often contain various geometric visual patterns, we utilize an object detection algorithm to ensure these patterns are not overlooked in the captioning process. We employed the SAM algorithm, which can detect various-size objects, to capture the visual features of these geometric patterns and used this information as input for the LLM. Under the puzzle split configuration, we achieved an option selection accuracy Oacc of 29.5 on the test set and a weighted option selection accuracy (WOSA) of 27.1 on the challenge set.
format Preprint
id arxiv_https___arxiv_org_abs_2406_05963
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Solution for SMART-101 Challenge of CVPR Multi-modal Algorithmic Reasoning Task 2024
Ahn, Jinwoo
Park, Junhyeok
Kim, Min-Jun
Kim, Kang-Hyeon
Sohn, So-Yeong
Lee, Yun-Ji
Chang, Du-Seong
Heo, Yu-Jung
Kim, Eun-Sol
Computer Vision and Pattern Recognition
Artificial Intelligence
In this paper, the solution of HYU MLLAB KT Team to the Multimodal Algorithmic Reasoning Task: SMART-101 CVPR 2024 Challenge is presented. Beyond conventional visual question-answering problems, the SMART-101 challenge aims to achieve human-level multimodal understanding by tackling complex visio-linguistic puzzles designed for children in the 6-8 age group. To solve this problem, we suggest two main ideas. First, to utilize the reasoning ability of a large-scale language model (LLM), the given visual cues (images) are grounded in the text modality. For this purpose, we generate highly detailed text captions that describe the context of the image and use these captions as input for the LLM. Second, due to the nature of puzzle images, which often contain various geometric visual patterns, we utilize an object detection algorithm to ensure these patterns are not overlooked in the captioning process. We employed the SAM algorithm, which can detect various-size objects, to capture the visual features of these geometric patterns and used this information as input for the LLM. Under the puzzle split configuration, we achieved an option selection accuracy Oacc of 29.5 on the test set and a weighted option selection accuracy (WOSA) of 27.1 on the challenge set.
title Solution for SMART-101 Challenge of CVPR Multi-modal Algorithmic Reasoning Task 2024
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2406.05963