Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Jiali, Hei, Xusen, Xue, Yuqi, Wei, Yuancheng, Xie, Jiayuan, Cai, Yi, Li, Qing
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913607388758016
author Chen, Jiali
Hei, Xusen
Xue, Yuqi
Wei, Yuancheng
Xie, Jiayuan
Cai, Yi
Li, Qing
author_facet Chen, Jiali
Hei, Xusen
Xue, Yuqi
Wei, Yuancheng
Xie, Jiayuan
Cai, Yi
Li, Qing
contents Large multimodal models (LMMs) have shown remarkable performance in the visual commonsense reasoning (VCR) task, which aims to answer a multiple-choice question based on visual commonsense within an image. However, the ability of LMMs to correct potential visual commonsense errors in the distractor upon their occurrence is yet under-explored. Drawing inspiration from how a human teacher crafts challenging distractors to test students' comprehension of the concepts or skills and assists them in identifying and correcting errors toward the answer, we are the pioneering research for LMMs to simulate this error correction process. To this end, we employ GPT-4 as a ``teacher'' to collect the explainable feedback dataset VCR-DF for error correction, which serves as a benchmark to evaluate the ability of LMMs to identify misconceptions and clarify reasons behind the error in VCR distractors toward final answers. In addition, we propose an LMM-based Pedagogical Expert Instructed Feedback Generation (PEIFG) model to incorporate the learnable expert prompts and multimodal instruction as guidance for feedback generation. Experimental results show that our PEIFG significantly outperforms existing LMMs. We believe that our benchmark provides a new direction for evaluating the capabilities of LMMs.
format Preprint
id arxiv_https___arxiv_org_abs_2412_07801
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor
Chen, Jiali
Hei, Xusen
Xue, Yuqi
Wei, Yuancheng
Xie, Jiayuan
Cai, Yi
Li, Qing
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Large multimodal models (LMMs) have shown remarkable performance in the visual commonsense reasoning (VCR) task, which aims to answer a multiple-choice question based on visual commonsense within an image. However, the ability of LMMs to correct potential visual commonsense errors in the distractor upon their occurrence is yet under-explored. Drawing inspiration from how a human teacher crafts challenging distractors to test students' comprehension of the concepts or skills and assists them in identifying and correcting errors toward the answer, we are the pioneering research for LMMs to simulate this error correction process. To this end, we employ GPT-4 as a ``teacher'' to collect the explainable feedback dataset VCR-DF for error correction, which serves as a benchmark to evaluate the ability of LMMs to identify misconceptions and clarify reasons behind the error in VCR distractors toward final answers. In addition, we propose an LMM-based Pedagogical Expert Instructed Feedback Generation (PEIFG) model to incorporate the learnable expert prompts and multimodal instruction as guidance for feedback generation. Experimental results show that our PEIFG significantly outperforms existing LMMs. We believe that our benchmark provides a new direction for evaluating the capabilities of LMMs.
title Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2412.07801