Saved in:
Bibliographic Details
Main Authors: Jiang, Jindong, Li, Xiuyu, Liu, Zhijian, Li, Muyang, Chen, Guo, Li, Zhiqi, Huang, De-An, Liu, Guilin, Yu, Zhiding, Keutzer, Kurt, Ahn, Sungjin, Kautz, Jan, Yin, Hongxu, Lu, Yao, Han, Song, Byeon, Wonmin
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2503.04130
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915507357089792
author Jiang, Jindong
Li, Xiuyu
Liu, Zhijian
Li, Muyang
Chen, Guo
Li, Zhiqi
Huang, De-An
Liu, Guilin
Yu, Zhiding
Keutzer, Kurt
Ahn, Sungjin
Kautz, Jan
Yin, Hongxu
Lu, Yao
Han, Song
Byeon, Wonmin
author_facet Jiang, Jindong
Li, Xiuyu
Liu, Zhijian
Li, Muyang
Chen, Guo
Li, Zhiqi
Huang, De-An
Liu, Guilin
Yu, Zhiding
Keutzer, Kurt
Ahn, Sungjin
Kautz, Jan
Yin, Hongxu
Lu, Yao
Han, Song
Byeon, Wonmin
contents Recent advances in video-based multimodal large language models (Video-LLMs) have significantly improved video understanding by processing videos as sequences of image frames. However, many existing methods treat frames independently in the vision backbone, lacking explicit temporal modeling, which limits their ability to capture dynamic patterns and efficiently handle long videos. To address these limitations, we introduce STORM (Spatiotemporal TOken Reduction for Multimodal LLMs), a novel architecture incorporating a dedicated temporal encoder between the image encoder and the LLM. Our temporal encoder leverages the Mamba State Space Model to integrate temporal information into image tokens, generating enriched representations that preserve inter-frame dynamics across the entire video sequence. This enriched encoding not only enhances video reasoning capabilities but also enables effective token reduction strategies, including test-time sampling and training-based temporal and spatial pooling, substantially reducing computational demands on the LLM without sacrificing key temporal information. By integrating these techniques, our approach simultaneously reduces training and inference latency while improving performance, enabling efficient and robust video understanding over extended temporal contexts. Extensive evaluations show that STORM achieves state-of-the-art results across various long video understanding benchmarks (more than 5% improvement on MLVU and LongVideoBench) while reducing the computation costs by up to $8\times$ and the decoding latency by 2.4-2.9$\times$ for the fixed numbers of input frames. Project page is available at https://research.nvidia.com/labs/lpr/storm
format Preprint
id arxiv_https___arxiv_org_abs_2503_04130
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle STORM: Token-Efficient Long Video Understanding for Multimodal LLMs
Jiang, Jindong
Li, Xiuyu
Liu, Zhijian
Li, Muyang
Chen, Guo
Li, Zhiqi
Huang, De-An
Liu, Guilin
Yu, Zhiding
Keutzer, Kurt
Ahn, Sungjin
Kautz, Jan
Yin, Hongxu
Lu, Yao
Han, Song
Byeon, Wonmin
Computer Vision and Pattern Recognition
Recent advances in video-based multimodal large language models (Video-LLMs) have significantly improved video understanding by processing videos as sequences of image frames. However, many existing methods treat frames independently in the vision backbone, lacking explicit temporal modeling, which limits their ability to capture dynamic patterns and efficiently handle long videos. To address these limitations, we introduce STORM (Spatiotemporal TOken Reduction for Multimodal LLMs), a novel architecture incorporating a dedicated temporal encoder between the image encoder and the LLM. Our temporal encoder leverages the Mamba State Space Model to integrate temporal information into image tokens, generating enriched representations that preserve inter-frame dynamics across the entire video sequence. This enriched encoding not only enhances video reasoning capabilities but also enables effective token reduction strategies, including test-time sampling and training-based temporal and spatial pooling, substantially reducing computational demands on the LLM without sacrificing key temporal information. By integrating these techniques, our approach simultaneously reduces training and inference latency while improving performance, enabling efficient and robust video understanding over extended temporal contexts. Extensive evaluations show that STORM achieves state-of-the-art results across various long video understanding benchmarks (more than 5% improvement on MLVU and LongVideoBench) while reducing the computation costs by up to $8\times$ and the decoding latency by 2.4-2.9$\times$ for the fixed numbers of input frames. Project page is available at https://research.nvidia.com/labs/lpr/storm
title STORM: Token-Efficient Long Video Understanding for Multimodal LLMs
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.04130