Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chatterjee, Dibyadip, Remelli, Edoardo, Song, Yale, Tekin, Bugra, Mittal, Abhay, Bhatnagar, Bharat, Camgöz, Necati Cihan, Hampali, Shreyas, Sauser, Eric, Ma, Shugao, Yao, Angela, Sener, Fadime
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908326440206336
author Chatterjee, Dibyadip
Remelli, Edoardo
Song, Yale
Tekin, Bugra
Mittal, Abhay
Bhatnagar, Bharat
Camgöz, Necati Cihan
Hampali, Shreyas
Sauser, Eric
Ma, Shugao
Yao, Angela
Sener, Fadime
author_facet Chatterjee, Dibyadip
Remelli, Edoardo
Song, Yale
Tekin, Bugra
Mittal, Abhay
Bhatnagar, Bharat
Camgöz, Necati Cihan
Hampali, Shreyas
Sauser, Eric
Ma, Shugao
Yao, Angela
Sener, Fadime
contents We introduce ProVideLLM, an end-to-end framework for real-time procedural video understanding. ProVideLLM integrates a multimodal cache configured to store two types of tokens - verbalized text tokens, which provide compressed textual summaries of long-term observations, and visual tokens, encoded with DETR-QFormer to capture fine-grained details from short-term observations. This design reduces token count by 22x over existing methods in representing one hour of long-term observations while effectively encoding fine-granularity of the present. By interleaving these tokens in our multimodal cache, ProVideLLM ensures sub-linear scaling of memory and compute with video length, enabling per-frame streaming inference at 10 FPS and streaming dialogue at 25 FPS, with a minimal 2GB GPU memory footprint. ProVideLLM also sets new state-of-the-art results on six procedural tasks across four datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2504_13915
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
Chatterjee, Dibyadip
Remelli, Edoardo
Song, Yale
Tekin, Bugra
Mittal, Abhay
Bhatnagar, Bharat
Camgöz, Necati Cihan
Hampali, Shreyas
Sauser, Eric
Ma, Shugao
Yao, Angela
Sener, Fadime
Computer Vision and Pattern Recognition
We introduce ProVideLLM, an end-to-end framework for real-time procedural video understanding. ProVideLLM integrates a multimodal cache configured to store two types of tokens - verbalized text tokens, which provide compressed textual summaries of long-term observations, and visual tokens, encoded with DETR-QFormer to capture fine-grained details from short-term observations. This design reduces token count by 22x over existing methods in representing one hour of long-term observations while effectively encoding fine-granularity of the present. By interleaving these tokens in our multimodal cache, ProVideLLM ensures sub-linear scaling of memory and compute with video length, enabling per-frame streaming inference at 10 FPS and streaming dialogue at 25 FPS, with a minimal 2GB GPU memory footprint. ProVideLLM also sets new state-of-the-art results on six procedural tasks across four datasets.
title Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.13915