MuDoC: An Interactive Multimodal Document-grounded Conversational AI System

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Taneja, Karan, Goel, Ashok K.
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929715325960192
author Taneja, Karan
Goel, Ashok K.
author_facet Taneja, Karan
Goel, Ashok K.
contents Multimodal AI is an important step towards building effective tools to leverage multiple modalities in human-AI communication. Building a multimodal document-grounded AI system to interact with long documents remains a challenge. Our work aims to fill the research gap of directly leveraging grounded visuals from documents alongside textual content in documents for response generation. We present an interactive conversational AI agent 'MuDoC' based on GPT-4o to generate document-grounded responses with interleaved text and figures. MuDoC's intelligent textbook interface promotes trustworthiness and enables verification of system responses by allowing instant navigation to source text and figures in the documents. We also discuss qualitative observations based on MuDoC responses highlighting its strengths and limitations.
format Preprint
id arxiv_https___arxiv_org_abs_2502_09843
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MuDoC: An Interactive Multimodal Document-grounded Conversational AI System
Taneja, Karan
Goel, Ashok K.
Artificial Intelligence
Human-Computer Interaction
Multimedia
Multimodal AI is an important step towards building effective tools to leverage multiple modalities in human-AI communication. Building a multimodal document-grounded AI system to interact with long documents remains a challenge. Our work aims to fill the research gap of directly leveraging grounded visuals from documents alongside textual content in documents for response generation. We present an interactive conversational AI agent 'MuDoC' based on GPT-4o to generate document-grounded responses with interleaved text and figures. MuDoC's intelligent textbook interface promotes trustworthiness and enables verification of system responses by allowing instant navigation to source text and figures in the documents. We also discuss qualitative observations based on MuDoC responses highlighting its strengths and limitations.
title MuDoC: An Interactive Multimodal Document-grounded Conversational AI System
topic Artificial Intelligence
Human-Computer Interaction
Multimedia
url https://arxiv.org/abs/2502.09843