Tracking Any Point with Frame-Event Fusion Network at High Frame Rate

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Jiaxiong, Wang, Bo, Tan, Zhen, Zhang, Jinpu, Shen, Hui, Hu, Dewen
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917779354943488
author Liu, Jiaxiong
Wang, Bo
Tan, Zhen
Zhang, Jinpu
Shen, Hui
Hu, Dewen
author_facet Liu, Jiaxiong
Wang, Bo
Tan, Zhen
Zhang, Jinpu
Shen, Hui
Hu, Dewen
contents Tracking any point based on image frames is constrained by frame rates, leading to instability in high-speed scenarios and limited generalization in real-world applications. To overcome these limitations, we propose an image-event fusion point tracker, FE-TAP, which combines the contextual information from image frames with the high temporal resolution of events, achieving high frame rate and robust point tracking under various challenging conditions. Specifically, we designed an Evolution Fusion module (EvoFusion) to model the image generation process guided by events. This module can effectively integrate valuable information from both modalities operating at different frequencies. To achieve smoother point trajectories, we employed a transformer-based refinement strategy that updates the point's trajectories and features iteratively. Extensive experiments demonstrate that our method outperforms state-of-the-art approaches, particularly improving expected feature age by 24$\%$ on EDS datasets. Finally, we qualitatively validated the robustness of our algorithm in real driving scenarios using our custom-designed high-resolution image-event synchronization device. Our source code will be released at https://github.com/ljx1002/FE-TAP.
format Preprint
id arxiv_https___arxiv_org_abs_2409_11953
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Tracking Any Point with Frame-Event Fusion Network at High Frame Rate
Liu, Jiaxiong
Wang, Bo
Tan, Zhen
Zhang, Jinpu
Shen, Hui
Hu, Dewen
Computer Vision and Pattern Recognition
Tracking any point based on image frames is constrained by frame rates, leading to instability in high-speed scenarios and limited generalization in real-world applications. To overcome these limitations, we propose an image-event fusion point tracker, FE-TAP, which combines the contextual information from image frames with the high temporal resolution of events, achieving high frame rate and robust point tracking under various challenging conditions. Specifically, we designed an Evolution Fusion module (EvoFusion) to model the image generation process guided by events. This module can effectively integrate valuable information from both modalities operating at different frequencies. To achieve smoother point trajectories, we employed a transformer-based refinement strategy that updates the point's trajectories and features iteratively. Extensive experiments demonstrate that our method outperforms state-of-the-art approaches, particularly improving expected feature age by 24$\%$ on EDS datasets. Finally, we qualitatively validated the robustness of our algorithm in real driving scenarios using our custom-designed high-resolution image-event synchronization device. Our source code will be released at https://github.com/ljx1002/FE-TAP.
title Tracking Any Point with Frame-Event Fusion Network at High Frame Rate
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2409.11953