Memory-QA: Answering Recall Questions Based on Multimodal Memories
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866911179007328256 |
|---|---|
| author | Jiang, Hongda Zhang, Xinyuan Garg, Siddhant Arora, Rishab Kuo, Shiun-Zu Xu, Jiayang Bansal, Ankur Brossman, Christopher Liu, Yue Colak, Aaron Aly, Ahmed Kumar, Anuj Dong, Xin Luna |
| author_facet | Jiang, Hongda Zhang, Xinyuan Garg, Siddhant Arora, Rishab Kuo, Shiun-Zu Xu, Jiayang Bansal, Ankur Brossman, Christopher Liu, Yue Colak, Aaron Aly, Ahmed Kumar, Anuj Dong, Xin Luna |
| contents | We introduce Memory-QA, a novel real-world task that involves answering recall questions about visual content from previously stored multimodal memories. This task poses unique challenges, including the creation of task-oriented memories, the effective utilization of temporal and location information within memories, and the ability to draw upon multiple memories to answer a recall question. To address these challenges, we propose a comprehensive pipeline, Pensieve, integrating memory-specific augmentation, time- and location-aware multi-signal retrieval, and multi-memory QA fine-tuning. We created a multimodal benchmark to illustrate various real challenges in this task, and show the superior performance of Pensieve over state-of-the-art solutions (up to 14% on QA accuracy). |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_18436 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Memory-QA: Answering Recall Questions Based on Multimodal Memories Jiang, Hongda Zhang, Xinyuan Garg, Siddhant Arora, Rishab Kuo, Shiun-Zu Xu, Jiayang Bansal, Ankur Brossman, Christopher Liu, Yue Colak, Aaron Aly, Ahmed Kumar, Anuj Dong, Xin Luna Artificial Intelligence Computation and Language Databases We introduce Memory-QA, a novel real-world task that involves answering recall questions about visual content from previously stored multimodal memories. This task poses unique challenges, including the creation of task-oriented memories, the effective utilization of temporal and location information within memories, and the ability to draw upon multiple memories to answer a recall question. To address these challenges, we propose a comprehensive pipeline, Pensieve, integrating memory-specific augmentation, time- and location-aware multi-signal retrieval, and multi-memory QA fine-tuning. We created a multimodal benchmark to illustrate various real challenges in this task, and show the superior performance of Pensieve over state-of-the-art solutions (up to 14% on QA accuracy). |
| title | Memory-QA: Answering Recall Questions Based on Multimodal Memories |
| topic | Artificial Intelligence Computation and Language Databases |
| url | https://arxiv.org/abs/2509.18436 |