Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision Encoders

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Rasekh, Ali, Soula, Erfan Bagheri, Daliran, Omid, Gottschalk, Simon, Fayyaz, Mohsen
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915586683961344
author Rasekh, Ali
Soula, Erfan Bagheri
Daliran, Omid
Gottschalk, Simon
Fayyaz, Mohsen
author_facet Rasekh, Ali
Soula, Erfan Bagheri
Daliran, Omid
Gottschalk, Simon
Fayyaz, Mohsen
contents Despite significant advances in Multimodal Large Language Models (MLLMs), understanding complex temporal dynamics in videos remains a major challenge. Our experiments show that current Video Large Language Model (Video-LLM) architectures have critical limitations in temporal understanding, struggling with tasks that require detailed comprehension of action sequences and temporal progression. In this work, we propose a Video-LLM architecture that introduces stacked temporal attention modules directly within the vision encoder. This design incorporates a temporal attention in vision encoder, enabling the model to better capture the progression of actions and the relationships between frames before passing visual tokens to the LLM. Our results show that this approach significantly improves temporal reasoning and outperforms existing models in video question answering tasks, specifically in action recognition. We improve on benchmarks including VITATECS, MVBench, and Video-MME by up to +5.5%. By enhancing the vision encoder with temporal structure, we address a critical gap in video understanding for Video-LLMs. Project page and code are available at: https://alirasekh.github.io/STAVEQ2/.
format Preprint
id arxiv_https___arxiv_org_abs_2510_26027
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision Encoders
Rasekh, Ali
Soula, Erfan Bagheri
Daliran, Omid
Gottschalk, Simon
Fayyaz, Mohsen
Computer Vision and Pattern Recognition
Despite significant advances in Multimodal Large Language Models (MLLMs), understanding complex temporal dynamics in videos remains a major challenge. Our experiments show that current Video Large Language Model (Video-LLM) architectures have critical limitations in temporal understanding, struggling with tasks that require detailed comprehension of action sequences and temporal progression. In this work, we propose a Video-LLM architecture that introduces stacked temporal attention modules directly within the vision encoder. This design incorporates a temporal attention in vision encoder, enabling the model to better capture the progression of actions and the relationships between frames before passing visual tokens to the LLM. Our results show that this approach significantly improves temporal reasoning and outperforms existing models in video question answering tasks, specifically in action recognition. We improve on benchmarks including VITATECS, MVBench, and Video-MME by up to +5.5%. By enhancing the vision encoder with temporal structure, we address a critical gap in video understanding for Video-LLMs. Project page and code are available at: https://alirasekh.github.io/STAVEQ2/.
title Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision Encoders
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.26027