OmniQuery: Contextually Augmenting Captured Multimodal Memory to Enable Personal Question Answering
Fuente:
arXiv
Salvato in:
| Autori principali: | , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866910837765046272 |
|---|---|
| author | Li, Jiahao Nick Zhang, Zhuohao Jerry Ma, Jiaju |
| author_facet | Li, Jiahao Nick Zhang, Zhuohao Jerry Ma, Jiaju |
| contents | People often capture memories through photos, screenshots, and videos. While existing AI-based tools enable querying this data using natural language, they only support retrieving individual pieces of information like certain objects in photos, and struggle with answering more complex queries that involve interpreting interconnected memories like sequential events. We conducted a one-month diary study to collect realistic user queries and generated a taxonomy of necessary contextual information for integrating with captured memories. We then introduce OmniQuery, a novel system that is able to answer complex personal memory-related questions that require extracting and inferring contextual information. OmniQuery augments individual captured memories through integrating scattered contextual information from multiple interconnected memories. Given a question, OmniQuery retrieves relevant augmented memories and uses a large language model (LLM) to generate answers with references. In human evaluations, we show the effectiveness of OmniQuery with an accuracy of 71.5%, outperforming a conventional RAG system by winning or tying for 74.5% of the time. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2409_08250 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | OmniQuery: Contextually Augmenting Captured Multimodal Memory to Enable Personal Question Answering Li, Jiahao Nick Zhang, Zhuohao Jerry Ma, Jiaju Human-Computer Interaction Artificial Intelligence People often capture memories through photos, screenshots, and videos. While existing AI-based tools enable querying this data using natural language, they only support retrieving individual pieces of information like certain objects in photos, and struggle with answering more complex queries that involve interpreting interconnected memories like sequential events. We conducted a one-month diary study to collect realistic user queries and generated a taxonomy of necessary contextual information for integrating with captured memories. We then introduce OmniQuery, a novel system that is able to answer complex personal memory-related questions that require extracting and inferring contextual information. OmniQuery augments individual captured memories through integrating scattered contextual information from multiple interconnected memories. Given a question, OmniQuery retrieves relevant augmented memories and uses a large language model (LLM) to generate answers with references. In human evaluations, we show the effectiveness of OmniQuery with an accuracy of 71.5%, outperforming a conventional RAG system by winning or tying for 74.5% of the time. |
| title | OmniQuery: Contextually Augmenting Captured Multimodal Memory to Enable Personal Question Answering |
| topic | Human-Computer Interaction Artificial Intelligence |
| url | https://arxiv.org/abs/2409.08250 |