Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866908326440206336 |
|---|---|
| author | Chatterjee, Dibyadip Remelli, Edoardo Song, Yale Tekin, Bugra Mittal, Abhay Bhatnagar, Bharat Camgöz, Necati Cihan Hampali, Shreyas Sauser, Eric Ma, Shugao Yao, Angela Sener, Fadime |
| author_facet | Chatterjee, Dibyadip Remelli, Edoardo Song, Yale Tekin, Bugra Mittal, Abhay Bhatnagar, Bharat Camgöz, Necati Cihan Hampali, Shreyas Sauser, Eric Ma, Shugao Yao, Angela Sener, Fadime |
| contents | We introduce ProVideLLM, an end-to-end framework for real-time procedural video understanding. ProVideLLM integrates a multimodal cache configured to store two types of tokens - verbalized text tokens, which provide compressed textual summaries of long-term observations, and visual tokens, encoded with DETR-QFormer to capture fine-grained details from short-term observations. This design reduces token count by 22x over existing methods in representing one hour of long-term observations while effectively encoding fine-granularity of the present. By interleaving these tokens in our multimodal cache, ProVideLLM ensures sub-linear scaling of memory and compute with video length, enabling per-frame streaming inference at 10 FPS and streaming dialogue at 25 FPS, with a minimal 2GB GPU memory footprint. ProVideLLM also sets new state-of-the-art results on six procedural tasks across four datasets. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_13915 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding Chatterjee, Dibyadip Remelli, Edoardo Song, Yale Tekin, Bugra Mittal, Abhay Bhatnagar, Bharat Camgöz, Necati Cihan Hampali, Shreyas Sauser, Eric Ma, Shugao Yao, Angela Sener, Fadime Computer Vision and Pattern Recognition We introduce ProVideLLM, an end-to-end framework for real-time procedural video understanding. ProVideLLM integrates a multimodal cache configured to store two types of tokens - verbalized text tokens, which provide compressed textual summaries of long-term observations, and visual tokens, encoded with DETR-QFormer to capture fine-grained details from short-term observations. This design reduces token count by 22x over existing methods in representing one hour of long-term observations while effectively encoding fine-granularity of the present. By interleaving these tokens in our multimodal cache, ProVideLLM ensures sub-linear scaling of memory and compute with video length, enabling per-frame streaming inference at 10 FPS and streaming dialogue at 25 FPS, with a minimal 2GB GPU memory footprint. ProVideLLM also sets new state-of-the-art results on six procedural tasks across four datasets. |
| title | Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2504.13915 |