Memory-QA: Answering Recall Questions Based on Multimodal Memories

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jiang, Hongda, Zhang, Xinyuan, Garg, Siddhant, Arora, Rishab, Kuo, Shiun-Zu, Xu, Jiayang, Bansal, Ankur, Brossman, Christopher, Liu, Yue, Colak, Aaron, Aly, Ahmed, Kumar, Anuj, Dong, Xin Luna
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911179007328256
author Jiang, Hongda
Zhang, Xinyuan
Garg, Siddhant
Arora, Rishab
Kuo, Shiun-Zu
Xu, Jiayang
Bansal, Ankur
Brossman, Christopher
Liu, Yue
Colak, Aaron
Aly, Ahmed
Kumar, Anuj
Dong, Xin Luna
author_facet Jiang, Hongda
Zhang, Xinyuan
Garg, Siddhant
Arora, Rishab
Kuo, Shiun-Zu
Xu, Jiayang
Bansal, Ankur
Brossman, Christopher
Liu, Yue
Colak, Aaron
Aly, Ahmed
Kumar, Anuj
Dong, Xin Luna
contents We introduce Memory-QA, a novel real-world task that involves answering recall questions about visual content from previously stored multimodal memories. This task poses unique challenges, including the creation of task-oriented memories, the effective utilization of temporal and location information within memories, and the ability to draw upon multiple memories to answer a recall question. To address these challenges, we propose a comprehensive pipeline, Pensieve, integrating memory-specific augmentation, time- and location-aware multi-signal retrieval, and multi-memory QA fine-tuning. We created a multimodal benchmark to illustrate various real challenges in this task, and show the superior performance of Pensieve over state-of-the-art solutions (up to 14% on QA accuracy).
format Preprint
id arxiv_https___arxiv_org_abs_2509_18436
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Memory-QA: Answering Recall Questions Based on Multimodal Memories
Jiang, Hongda
Zhang, Xinyuan
Garg, Siddhant
Arora, Rishab
Kuo, Shiun-Zu
Xu, Jiayang
Bansal, Ankur
Brossman, Christopher
Liu, Yue
Colak, Aaron
Aly, Ahmed
Kumar, Anuj
Dong, Xin Luna
Artificial Intelligence
Computation and Language
Databases
We introduce Memory-QA, a novel real-world task that involves answering recall questions about visual content from previously stored multimodal memories. This task poses unique challenges, including the creation of task-oriented memories, the effective utilization of temporal and location information within memories, and the ability to draw upon multiple memories to answer a recall question. To address these challenges, we propose a comprehensive pipeline, Pensieve, integrating memory-specific augmentation, time- and location-aware multi-signal retrieval, and multi-memory QA fine-tuning. We created a multimodal benchmark to illustrate various real challenges in this task, and show the superior performance of Pensieve over state-of-the-art solutions (up to 14% on QA accuracy).
title Memory-QA: Answering Recall Questions Based on Multimodal Memories
topic Artificial Intelligence
Computation and Language
Databases
url https://arxiv.org/abs/2509.18436