Generative Event Pretraining with Foundation Model Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cao, Jianwen, Xing, Jiaxu, Messikommer, Nico, Scaramuzza, Davide
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915911468843008
author Cao, Jianwen
Xing, Jiaxu
Messikommer, Nico
Scaramuzza, Davide
author_facet Cao, Jianwen
Xing, Jiaxu
Messikommer, Nico
Scaramuzza, Davide
contents Event cameras provide robust visual signals under fast motion and challenging illumination conditions thanks to their microsecond latency and high dynamic range. However, their unique sensing characteristics and limited labeled data make it challenging to train event-based visual foundation models (VFMs), which are crucial for learning visual features transferable across tasks. To tackle this problem, we propose GEP (Generative Event Pretraining), a two-stage framework that transfers semantic knowledge learned from internet-scale image datasets to event data while learning event-specific temporal dynamics. First, an event encoder is aligned to a frozen VFM through a joint regression-contrastive objective, grounding event features in image semantics. Second, a transformer backbone is autoregressively pretrained on mixed event-image sequences to capture the temporal structure unique to events. Our approach outperforms state-of-the-art event pretraining methods on a diverse range of downstream tasks, including object recognition, segmentation, and depth estimation. Together, VFM-guided alignment and generative sequence modeling yield a semantically rich, temporally aware event model that generalizes robustly across domains.
format Preprint
id arxiv_https___arxiv_org_abs_2603_23032
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Generative Event Pretraining with Foundation Model Alignment
Cao, Jianwen
Xing, Jiaxu
Messikommer, Nico
Scaramuzza, Davide
Computer Vision and Pattern Recognition
Robotics
Event cameras provide robust visual signals under fast motion and challenging illumination conditions thanks to their microsecond latency and high dynamic range. However, their unique sensing characteristics and limited labeled data make it challenging to train event-based visual foundation models (VFMs), which are crucial for learning visual features transferable across tasks. To tackle this problem, we propose GEP (Generative Event Pretraining), a two-stage framework that transfers semantic knowledge learned from internet-scale image datasets to event data while learning event-specific temporal dynamics. First, an event encoder is aligned to a frozen VFM through a joint regression-contrastive objective, grounding event features in image semantics. Second, a transformer backbone is autoregressively pretrained on mixed event-image sequences to capture the temporal structure unique to events. Our approach outperforms state-of-the-art event pretraining methods on a diverse range of downstream tasks, including object recognition, segmentation, and depth estimation. Together, VFM-guided alignment and generative sequence modeling yield a semantically rich, temporally aware event model that generalizes robustly across domains.
title Generative Event Pretraining with Foundation Model Alignment
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2603.23032