AUGUSTUS: An LLM-Driven Multimodal Agent System with Contextualized User Memory

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Jain, Jitesh, Maheshwari, Shubham, Yu, Ning, Hwu, Wen-mei, Shi, Humphrey
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917021183115264
author Jain, Jitesh
Maheshwari, Shubham
Yu, Ning
Hwu, Wen-mei
Shi, Humphrey
author_facet Jain, Jitesh
Maheshwari, Shubham
Yu, Ning
Hwu, Wen-mei
Shi, Humphrey
contents Riding on the success of LLMs with retrieval-augmented generation (RAG), there has been a growing interest in augmenting agent systems with external memory databases. However, the existing systems focus on storing text information in their memory, ignoring the importance of multimodal signals. Motivated by the multimodal nature of human memory, we present AUGUSTUS, a multimodal agent system aligned with the ideas of human memory in cognitive science. Technically, our system consists of 4 stages connected in a loop: (i) encode: understanding the inputs; (ii) store in memory: saving important information; (iii) retrieve: searching for relevant context from memory; and (iv) act: perform the task. Unlike existing systems that use vector databases, we propose conceptualizing information into semantic tags and associating the tags with their context to store them in a graph-structured multimodal contextual memory for efficient concept-driven retrieval. Our system outperforms the traditional multimodal RAG approach while being 3.5 times faster for ImageNet classification and outperforming MemGPT on the MSC benchmark.
format Preprint
id arxiv_https___arxiv_org_abs_2510_15261
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AUGUSTUS: An LLM-Driven Multimodal Agent System with Contextualized User Memory
Jain, Jitesh
Maheshwari, Shubham
Yu, Ning
Hwu, Wen-mei
Shi, Humphrey
Artificial Intelligence
Riding on the success of LLMs with retrieval-augmented generation (RAG), there has been a growing interest in augmenting agent systems with external memory databases. However, the existing systems focus on storing text information in their memory, ignoring the importance of multimodal signals. Motivated by the multimodal nature of human memory, we present AUGUSTUS, a multimodal agent system aligned with the ideas of human memory in cognitive science. Technically, our system consists of 4 stages connected in a loop: (i) encode: understanding the inputs; (ii) store in memory: saving important information; (iii) retrieve: searching for relevant context from memory; and (iv) act: perform the task. Unlike existing systems that use vector databases, we propose conceptualizing information into semantic tags and associating the tags with their context to store them in a graph-structured multimodal contextual memory for efficient concept-driven retrieval. Our system outperforms the traditional multimodal RAG approach while being 3.5 times faster for ImageNet classification and outperforming MemGPT on the MSC benchmark.
title AUGUSTUS: An LLM-Driven Multimodal Agent System with Contextualized User Memory
topic Artificial Intelligence
url https://arxiv.org/abs/2510.15261