Memory-Augmented Multimodal LLMs for Surgical VQA via Self-Contained Inquiry

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Hou, Wenjun, Cheng, Yi, Xu, Kaishuai, Hu, Yan, Li, Wenjie, Liu, Jiang
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910702379204608
author Hou, Wenjun
Cheng, Yi
Xu, Kaishuai
Hu, Yan
Li, Wenjie
Liu, Jiang
author_facet Hou, Wenjun
Cheng, Yi
Xu, Kaishuai
Hu, Yan
Li, Wenjie
Liu, Jiang
contents Comprehensively understanding surgical scenes in Surgical Visual Question Answering (Surgical VQA) requires reasoning over multiple objects. Previous approaches address this task using cross-modal fusion strategies to enhance reasoning ability. However, these methods often struggle with limited scene understanding and question comprehension, and some rely on external resources (e.g., pre-extracted object features), which can introduce errors and generalize poorly across diverse surgical environments. To address these challenges, we propose SCAN, a simple yet effective memory-augmented framework that leverages Multimodal LLMs to improve surgical context comprehension via Self-Contained Inquiry. SCAN operates autonomously, generating two types of memory for context augmentation: Direct Memory (DM), which provides multiple candidates (or hints) to the final answer, and Indirect Memory (IM), which consists of self-contained question-hint pairs to capture broader scene context. DM directly assists in answering the question, while IM enhances understanding of the surgical scene beyond the immediate query. Reasoning over these object-aware memories enables the model to accurately interpret images and respond to questions. Extensive experiments on three publicly available Surgical VQA datasets demonstrate that SCAN achieves state-of-the-art performance, offering improved accuracy and robustness across various surgical scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2411_10937
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Memory-Augmented Multimodal LLMs for Surgical VQA via Self-Contained Inquiry
Hou, Wenjun
Cheng, Yi
Xu, Kaishuai
Hu, Yan
Li, Wenjie
Liu, Jiang
Computer Vision and Pattern Recognition
Computation and Language
Comprehensively understanding surgical scenes in Surgical Visual Question Answering (Surgical VQA) requires reasoning over multiple objects. Previous approaches address this task using cross-modal fusion strategies to enhance reasoning ability. However, these methods often struggle with limited scene understanding and question comprehension, and some rely on external resources (e.g., pre-extracted object features), which can introduce errors and generalize poorly across diverse surgical environments. To address these challenges, we propose SCAN, a simple yet effective memory-augmented framework that leverages Multimodal LLMs to improve surgical context comprehension via Self-Contained Inquiry. SCAN operates autonomously, generating two types of memory for context augmentation: Direct Memory (DM), which provides multiple candidates (or hints) to the final answer, and Indirect Memory (IM), which consists of self-contained question-hint pairs to capture broader scene context. DM directly assists in answering the question, while IM enhances understanding of the surgical scene beyond the immediate query. Reasoning over these object-aware memories enables the model to accurately interpret images and respond to questions. Extensive experiments on three publicly available Surgical VQA datasets demonstrate that SCAN achieves state-of-the-art performance, offering improved accuracy and robustness across various surgical scenarios.
title Memory-Augmented Multimodal LLMs for Surgical VQA via Self-Contained Inquiry
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2411.10937