HiCM$^2$: Hierarchical Compact Memory Modeling for Dense Video Captioning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kim, Minkuk, Kim, Hyeon Bae, Moon, Jinyoung, Choi, Jinwoo, Kim, Seong Tae
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917873850515456
author Kim, Minkuk
Kim, Hyeon Bae
Moon, Jinyoung
Choi, Jinwoo
Kim, Seong Tae
author_facet Kim, Minkuk
Kim, Hyeon Bae
Moon, Jinyoung
Choi, Jinwoo
Kim, Seong Tae
contents With the growing demand for solutions to real-world video challenges, interest in dense video captioning (DVC) has been on the rise. DVC involves the automatic captioning and localization of untrimmed videos. Several studies highlight the challenges of DVC and introduce improved methods utilizing prior knowledge, such as pre-training and external memory. In this research, we propose a model that leverages the prior knowledge of human-oriented hierarchical compact memory inspired by human memory hierarchy and cognition. To mimic human-like memory recall, we construct a hierarchical memory and a hierarchical memory reading module. We build an efficient hierarchical compact memory by employing clustering of memory events and summarization using large language models. Comparative experiments demonstrate that this hierarchical memory recall process improves the performance of DVC by achieving state-of-the-art performance on YouCook2 and ViTT datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2412_14585
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle HiCM$^2$: Hierarchical Compact Memory Modeling for Dense Video Captioning
Kim, Minkuk
Kim, Hyeon Bae
Moon, Jinyoung
Choi, Jinwoo
Kim, Seong Tae
Computer Vision and Pattern Recognition
With the growing demand for solutions to real-world video challenges, interest in dense video captioning (DVC) has been on the rise. DVC involves the automatic captioning and localization of untrimmed videos. Several studies highlight the challenges of DVC and introduce improved methods utilizing prior knowledge, such as pre-training and external memory. In this research, we propose a model that leverages the prior knowledge of human-oriented hierarchical compact memory inspired by human memory hierarchy and cognition. To mimic human-like memory recall, we construct a hierarchical memory and a hierarchical memory reading module. We build an efficient hierarchical compact memory by employing clustering of memory events and summarization using large language models. Comparative experiments demonstrate that this hierarchical memory recall process improves the performance of DVC by achieving state-of-the-art performance on YouCook2 and ViTT datasets.
title HiCM$^2$: Hierarchical Compact Memory Modeling for Dense Video Captioning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.14585