HiCM$^2$: Hierarchical Compact Memory Modeling for Dense Video Captioning
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866917873850515456 |
|---|---|
| author | Kim, Minkuk Kim, Hyeon Bae Moon, Jinyoung Choi, Jinwoo Kim, Seong Tae |
| author_facet | Kim, Minkuk Kim, Hyeon Bae Moon, Jinyoung Choi, Jinwoo Kim, Seong Tae |
| contents | With the growing demand for solutions to real-world video challenges, interest in dense video captioning (DVC) has been on the rise. DVC involves the automatic captioning and localization of untrimmed videos. Several studies highlight the challenges of DVC and introduce improved methods utilizing prior knowledge, such as pre-training and external memory. In this research, we propose a model that leverages the prior knowledge of human-oriented hierarchical compact memory inspired by human memory hierarchy and cognition. To mimic human-like memory recall, we construct a hierarchical memory and a hierarchical memory reading module. We build an efficient hierarchical compact memory by employing clustering of memory events and summarization using large language models. Comparative experiments demonstrate that this hierarchical memory recall process improves the performance of DVC by achieving state-of-the-art performance on YouCook2 and ViTT datasets. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_14585 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | HiCM$^2$: Hierarchical Compact Memory Modeling for Dense Video Captioning Kim, Minkuk Kim, Hyeon Bae Moon, Jinyoung Choi, Jinwoo Kim, Seong Tae Computer Vision and Pattern Recognition With the growing demand for solutions to real-world video challenges, interest in dense video captioning (DVC) has been on the rise. DVC involves the automatic captioning and localization of untrimmed videos. Several studies highlight the challenges of DVC and introduce improved methods utilizing prior knowledge, such as pre-training and external memory. In this research, we propose a model that leverages the prior knowledge of human-oriented hierarchical compact memory inspired by human memory hierarchy and cognition. To mimic human-like memory recall, we construct a hierarchical memory and a hierarchical memory reading module. We build an efficient hierarchical compact memory by employing clustering of memory events and summarization using large language models. Comparative experiments demonstrate that this hierarchical memory recall process improves the performance of DVC by achieving state-of-the-art performance on YouCook2 and ViTT datasets. |
| title | HiCM$^2$: Hierarchical Compact Memory Modeling for Dense Video Captioning |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2412.14585 |