Second Place Solution of WSDM2023 Toloka Visual Question Answering Challenge

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wu, Xiangyu, Chi, Zhouyang, Yang, Yang, Lu, Jianfeng
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917713516953600
author Wu, Xiangyu
Chi, Zhouyang
Yang, Yang
Lu, Jianfeng
author_facet Wu, Xiangyu
Chi, Zhouyang
Yang, Yang
Lu, Jianfeng
contents In this paper, we present our solution for the WSDM2023 Toloka Visual Question Answering Challenge. Inspired by the application of multimodal pre-trained models to various downstream tasks(e.g., visual question answering, visual grounding, and cross-modal retrieval), we approached this competition as a visual grounding task, where the input is an image and a question, guiding the model to answer the question and display the answer as a bounding box on the image. We designed a three-stage solution for this task. Specifically, we used the visual-language pre-trained model OFA as the foundation. In the first stage, we constructed a large-scale synthetic dataset similar to the competition dataset and coarse-tuned the model to learn generalized semantic information. In the second stage, we treated the competition task as a visual grounding task, loaded the weights from the previous stage, and continued to fine-tune the model on the competition dataset, transferring the semantic information learned in the first stage to the competition task. Finally, we designed a bounding box matching and replacing post-processing strategy to correct the model's prediction results. Our team achieved a score of 76.342 on the final leaderboard, ranking second.
format Preprint
id arxiv_https___arxiv_org_abs_2407_04255
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Second Place Solution of WSDM2023 Toloka Visual Question Answering Challenge
Wu, Xiangyu
Chi, Zhouyang
Yang, Yang
Lu, Jianfeng
Computer Vision and Pattern Recognition
In this paper, we present our solution for the WSDM2023 Toloka Visual Question Answering Challenge. Inspired by the application of multimodal pre-trained models to various downstream tasks(e.g., visual question answering, visual grounding, and cross-modal retrieval), we approached this competition as a visual grounding task, where the input is an image and a question, guiding the model to answer the question and display the answer as a bounding box on the image. We designed a three-stage solution for this task. Specifically, we used the visual-language pre-trained model OFA as the foundation. In the first stage, we constructed a large-scale synthetic dataset similar to the competition dataset and coarse-tuned the model to learn generalized semantic information. In the second stage, we treated the competition task as a visual grounding task, loaded the weights from the previous stage, and continued to fine-tune the model on the competition dataset, transferring the semantic information learned in the first stage to the competition task. Finally, we designed a bounding box matching and replacing post-processing strategy to correct the model's prediction results. Our team achieved a score of 76.342 on the final leaderboard, ranking second.
title Second Place Solution of WSDM2023 Toloka Visual Question Answering Challenge
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.04255