Temporal-Guided Visual Foundation Models for Event-Based Vision

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xia, Ruihao, Cai, Junhong, Leng, Luziwei, Wang, Liuyi, Liu, Chengju, Cheng, Ran, Tang, Yang, Zhou, Pan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912696068210688
author Xia, Ruihao
Cai, Junhong
Leng, Luziwei
Wang, Liuyi
Liu, Chengju
Cheng, Ran
Tang, Yang
Zhou, Pan
author_facet Xia, Ruihao
Cai, Junhong
Leng, Luziwei
Wang, Liuyi
Liu, Chengju
Cheng, Ran
Tang, Yang
Zhou, Pan
contents Event cameras offer unique advantages for vision tasks in challenging environments, yet processing asynchronous event streams remains an open challenge. While existing methods rely on specialized architectures or resource-intensive training, the potential of leveraging modern Visual Foundation Models (VFMs) pretrained on image data remains under-explored for event-based vision. To address this, we propose Temporal-Guided VFM (TGVFM), a novel framework that integrates VFMs with our temporal context fusion block seamlessly to bridge this gap. Our temporal block introduces three key components: (1) Long-Range Temporal Attention to model global temporal dependencies, (2) Dual Spatiotemporal Attention for multi-scale frame correlation, and (3) Deep Feature Guidance Mechanism to fuse semantic-temporal features. By retraining event-to-video models on real-world data and leveraging transformer-based VFMs, TGVFM preserves spatiotemporal dynamics while harnessing pretrained representations. Experiments demonstrate SoTA performance across semantic segmentation, depth estimation, and object detection, with improvements of 16%, 21%, and 16% over existing methods, respectively. Overall, this work unlocks the cross-modality potential of image-based VFMs for event-based vision with temporal reasoning. Code is available at https://github.com/XiaRho/TGVFM.
format Preprint
id arxiv_https___arxiv_org_abs_2511_06238
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Temporal-Guided Visual Foundation Models for Event-Based Vision
Xia, Ruihao
Cai, Junhong
Leng, Luziwei
Wang, Liuyi
Liu, Chengju
Cheng, Ran
Tang, Yang
Zhou, Pan
Computer Vision and Pattern Recognition
Event cameras offer unique advantages for vision tasks in challenging environments, yet processing asynchronous event streams remains an open challenge. While existing methods rely on specialized architectures or resource-intensive training, the potential of leveraging modern Visual Foundation Models (VFMs) pretrained on image data remains under-explored for event-based vision. To address this, we propose Temporal-Guided VFM (TGVFM), a novel framework that integrates VFMs with our temporal context fusion block seamlessly to bridge this gap. Our temporal block introduces three key components: (1) Long-Range Temporal Attention to model global temporal dependencies, (2) Dual Spatiotemporal Attention for multi-scale frame correlation, and (3) Deep Feature Guidance Mechanism to fuse semantic-temporal features. By retraining event-to-video models on real-world data and leveraging transformer-based VFMs, TGVFM preserves spatiotemporal dynamics while harnessing pretrained representations. Experiments demonstrate SoTA performance across semantic segmentation, depth estimation, and object detection, with improvements of 16%, 21%, and 16% over existing methods, respectively. Overall, this work unlocks the cross-modality potential of image-based VFMs for event-based vision with temporal reasoning. Code is available at https://github.com/XiaRho/TGVFM.
title Temporal-Guided Visual Foundation Models for Event-Based Vision
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.06238