Towards a Multimodal Document-grounded Conversational AI System for Education

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Taneja, Karan, Singh, Anjali, Goel, Ashok K.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912334807564288
author Taneja, Karan
Singh, Anjali
Goel, Ashok K.
author_facet Taneja, Karan
Singh, Anjali
Goel, Ashok K.
contents Multimedia learning using text and images has been shown to improve learning outcomes compared to text-only instruction. But conversational AI systems in education predominantly rely on text-based interactions while multimodal conversations for multimedia learning remain unexplored. Moreover, deploying conversational AI in learning contexts requires grounding in reliable sources and verifiability to create trust. We present MuDoC, a Multimodal Document-grounded Conversational AI system based on GPT-4o, that leverages both text and visuals from documents to generate responses interleaved with text and images. Its interface allows verification of AI generated content through seamless navigation to the source. We compare MuDoC to a text-only system to explore differences in learner engagement, trust in AI system, and their performance on problem-solving tasks. Our findings indicate that both visuals and verifiability of content enhance learner engagement and foster trust; however, no significant impact in performance was observed. We draw upon theories from cognitive and learning sciences to interpret the findings and derive implications, and outline future directions for the development of multimodal conversational AI systems in education.
format Preprint
id arxiv_https___arxiv_org_abs_2504_13884
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards a Multimodal Document-grounded Conversational AI System for Education
Taneja, Karan
Singh, Anjali
Goel, Ashok K.
Human-Computer Interaction
Artificial Intelligence
Computer Vision and Pattern Recognition
Multimedia learning using text and images has been shown to improve learning outcomes compared to text-only instruction. But conversational AI systems in education predominantly rely on text-based interactions while multimodal conversations for multimedia learning remain unexplored. Moreover, deploying conversational AI in learning contexts requires grounding in reliable sources and verifiability to create trust. We present MuDoC, a Multimodal Document-grounded Conversational AI system based on GPT-4o, that leverages both text and visuals from documents to generate responses interleaved with text and images. Its interface allows verification of AI generated content through seamless navigation to the source. We compare MuDoC to a text-only system to explore differences in learner engagement, trust in AI system, and their performance on problem-solving tasks. Our findings indicate that both visuals and verifiability of content enhance learner engagement and foster trust; however, no significant impact in performance was observed. We draw upon theories from cognitive and learning sciences to interpret the findings and derive implications, and outline future directions for the development of multimodal conversational AI systems in education.
title Towards a Multimodal Document-grounded Conversational AI System for Education
topic Human-Computer Interaction
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.13884