Beyond the Textual: Generating Coherent Visual Options for MCQs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Wanqiang, He, Longzhu, Zheng, Wei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916918131163136
author Wang, Wanqiang
He, Longzhu
Zheng, Wei
author_facet Wang, Wanqiang
He, Longzhu
Zheng, Wei
contents Multiple-choice questions (MCQs) play a crucial role in fostering deep thinking and knowledge integration in education. However, previous research has primarily focused on generating MCQs with textual options, but it largely overlooks the visual options. Moreover, generating high-quality distractors remains a major challenge due to the high cost and limited scalability of manual authoring. To tackle these problems, we propose a Cross-modal Options Synthesis (CmOS), a novel framework for generating educational MCQs with visual options. Our framework integrates Multimodal Chain-of-Thought (MCoT) reasoning process and Retrieval-Augmented Generation (RAG) to produce semantically plausible and visually similar answer and distractors. It also includes a discrimination module to identify content suitable for visual options. Experimental results on test tasks demonstrate the superiority of CmOS in content discrimination, question generation and visual option generation over existing methods across various subjects and educational levels.
format Preprint
id arxiv_https___arxiv_org_abs_2508_18772
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond the Textual: Generating Coherent Visual Options for MCQs
Wang, Wanqiang
He, Longzhu
Zheng, Wei
Computer Vision and Pattern Recognition
Computation and Language
Multiple-choice questions (MCQs) play a crucial role in fostering deep thinking and knowledge integration in education. However, previous research has primarily focused on generating MCQs with textual options, but it largely overlooks the visual options. Moreover, generating high-quality distractors remains a major challenge due to the high cost and limited scalability of manual authoring. To tackle these problems, we propose a Cross-modal Options Synthesis (CmOS), a novel framework for generating educational MCQs with visual options. Our framework integrates Multimodal Chain-of-Thought (MCoT) reasoning process and Retrieval-Augmented Generation (RAG) to produce semantically plausible and visually similar answer and distractors. It also includes a discrimination module to identify content suitable for visual options. Experimental results on test tasks demonstrate the superiority of CmOS in content discrimination, question generation and visual option generation over existing methods across various subjects and educational levels.
title Beyond the Textual: Generating Coherent Visual Options for MCQs
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2508.18772