Beyond Retrieval: Joint Supervision and Multimodal Document Ranking for Textbook Question Answering

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Alawwad, Hessa, Naseem, Usman, Alhothali, Areej, Alkhathlan, Ali, Jamal, Amani
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918025537519616
author Alawwad, Hessa
Naseem, Usman
Alhothali, Areej
Alkhathlan, Ali
Jamal, Amani
author_facet Alawwad, Hessa
Naseem, Usman
Alhothali, Areej
Alkhathlan, Ali
Jamal, Amani
contents Textbook question answering (TQA) is a complex task, requiring the interpretation of complex multimodal context. Although recent advances have improved overall performance, they often encounter difficulties in educational settings where accurate semantic alignment and task-specific document retrieval are essential. In this paper, we propose a novel approach to multimodal textbook question answering by introducing a mechanism for enhancing semantic representations through multi-objective joint training. Our model, Joint Embedding Training With Ranking Supervision for Textbook Question Answering (JETRTQA), is a multimodal learning framework built on a retriever--generator architecture that uses a retrieval-augmented generation setup, in which a multimodal large language model generates answers. JETRTQA is designed to improve the relevance of retrieved documents in complex educational contexts. Unlike traditional direct scoring approaches, JETRTQA learns to refine the semantic representations of questions and documents through a supervised signal that combines pairwise ranking and implicit supervision derived from answers. We evaluate our method on the CK12-QA dataset and demonstrate that it significantly improves the discrimination between informative and irrelevant documents, even when they are long, complex, and multimodal. JETRTQA outperforms the previous state of the art, achieving a 2.4\% gain in accuracy on the validation set and 11.1\% on the test set.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13520
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Retrieval: Joint Supervision and Multimodal Document Ranking for Textbook Question Answering
Alawwad, Hessa
Naseem, Usman
Alhothali, Areej
Alkhathlan, Ali
Jamal, Amani
Information Retrieval
Artificial Intelligence
Textbook question answering (TQA) is a complex task, requiring the interpretation of complex multimodal context. Although recent advances have improved overall performance, they often encounter difficulties in educational settings where accurate semantic alignment and task-specific document retrieval are essential. In this paper, we propose a novel approach to multimodal textbook question answering by introducing a mechanism for enhancing semantic representations through multi-objective joint training. Our model, Joint Embedding Training With Ranking Supervision for Textbook Question Answering (JETRTQA), is a multimodal learning framework built on a retriever--generator architecture that uses a retrieval-augmented generation setup, in which a multimodal large language model generates answers. JETRTQA is designed to improve the relevance of retrieved documents in complex educational contexts. Unlike traditional direct scoring approaches, JETRTQA learns to refine the semantic representations of questions and documents through a supervised signal that combines pairwise ranking and implicit supervision derived from answers. We evaluate our method on the CK12-QA dataset and demonstrate that it significantly improves the discrimination between informative and irrelevant documents, even when they are long, complex, and multimodal. JETRTQA outperforms the previous state of the art, achieving a 2.4\% gain in accuracy on the validation set and 11.1\% on the test set.
title Beyond Retrieval: Joint Supervision and Multimodal Document Ranking for Textbook Question Answering
topic Information Retrieval
Artificial Intelligence
url https://arxiv.org/abs/2505.13520