MuRAR: A Simple and Effective Multimodal Retrieval and Answer Refinement Framework for Multimodal Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Zhengyuan, Lee, Daniel, Zhang, Hong, Harsha, Sai Sree, Feujio, Loic, Maharaj, Akash, Li, Yunyao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916601926778880
author Zhu, Zhengyuan
Lee, Daniel
Zhang, Hong
Harsha, Sai Sree
Feujio, Loic
Maharaj, Akash
Li, Yunyao
author_facet Zhu, Zhengyuan
Lee, Daniel
Zhang, Hong
Harsha, Sai Sree
Feujio, Loic
Maharaj, Akash
Li, Yunyao
contents Recent advancements in retrieval-augmented generation (RAG) have demonstrated impressive performance in the question-answering (QA) task. However, most previous works predominantly focus on text-based answers. While some studies address multimodal data, they still fall short in generating comprehensive multimodal answers, particularly for explaining concepts or providing step-by-step tutorials on how to accomplish specific goals. This capability is especially valuable for applications such as enterprise chatbots and settings such as customer service and educational systems, where the answers are sourced from multimodal data. In this paper, we introduce a simple and effective framework named MuRAR (Multimodal Retrieval and Answer Refinement). MuRAR enhances text-based answers by retrieving relevant multimodal data and refining the responses to create coherent multimodal answers. This framework can be easily extended to support multimodal answers in enterprise chatbots with minimal modifications. Human evaluation results indicate that multimodal answers generated by MuRAR are more useful and readable compared to plain text answers.
format Preprint
id arxiv_https___arxiv_org_abs_2408_08521
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MuRAR: A Simple and Effective Multimodal Retrieval and Answer Refinement Framework for Multimodal Question Answering
Zhu, Zhengyuan
Lee, Daniel
Zhang, Hong
Harsha, Sai Sree
Feujio, Loic
Maharaj, Akash
Li, Yunyao
Information Retrieval
Computation and Language
Recent advancements in retrieval-augmented generation (RAG) have demonstrated impressive performance in the question-answering (QA) task. However, most previous works predominantly focus on text-based answers. While some studies address multimodal data, they still fall short in generating comprehensive multimodal answers, particularly for explaining concepts or providing step-by-step tutorials on how to accomplish specific goals. This capability is especially valuable for applications such as enterprise chatbots and settings such as customer service and educational systems, where the answers are sourced from multimodal data. In this paper, we introduce a simple and effective framework named MuRAR (Multimodal Retrieval and Answer Refinement). MuRAR enhances text-based answers by retrieving relevant multimodal data and refining the responses to create coherent multimodal answers. This framework can be easily extended to support multimodal answers in enterprise chatbots with minimal modifications. Human evaluation results indicate that multimodal answers generated by MuRAR are more useful and readable compared to plain text answers.
title MuRAR: A Simple and Effective Multimodal Retrieval and Answer Refinement Framework for Multimodal Question Answering
topic Information Retrieval
Computation and Language
url https://arxiv.org/abs/2408.08521