SOTFormer: A Minimal Transformer for Unified Object Tracking and Trajectory Prediction

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Dong, Zhongping, Yu, Pengyang, Li, Shuangjian, Chen, Liming, Kechadi, Mohand Tahar
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917081028493312
author Dong, Zhongping
Yu, Pengyang
Li, Shuangjian
Chen, Liming
Kechadi, Mohand Tahar
author_facet Dong, Zhongping
Yu, Pengyang
Li, Shuangjian
Chen, Liming
Kechadi, Mohand Tahar
contents Accurate single-object tracking and short-term motion forecasting remain challenging under occlusion, scale variation, and temporal drift, which disrupt the temporal coherence required for real-time perception. We introduce \textbf{SOTFormer}, a minimal constant-memory temporal transformer that unifies object detection, tracking, and short-horizon trajectory prediction within a single end-to-end framework. Unlike prior models with recurrent or stacked temporal encoders, SOTFormer achieves stable identity propagation through a ground-truth-primed memory and a burn-in anchor loss that explicitly stabilizes initialization. A single lightweight temporal-attention layer refines embeddings across frames, enabling real-time inference with fixed GPU memory. On the Mini-LaSOT (20%) benchmark, SOTFormer attains 76.3 AUC and 53.7 FPS (AMP, 4.3 GB VRAM), outperforming transformer baselines such as TrackFormer and MOTRv2 under fast motion, scale change, and occlusion.
format Preprint
id arxiv_https___arxiv_org_abs_2511_11824
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SOTFormer: A Minimal Transformer for Unified Object Tracking and Trajectory Prediction
Dong, Zhongping
Yu, Pengyang
Li, Shuangjian
Chen, Liming
Kechadi, Mohand Tahar
Computer Vision and Pattern Recognition
Accurate single-object tracking and short-term motion forecasting remain challenging under occlusion, scale variation, and temporal drift, which disrupt the temporal coherence required for real-time perception. We introduce \textbf{SOTFormer}, a minimal constant-memory temporal transformer that unifies object detection, tracking, and short-horizon trajectory prediction within a single end-to-end framework. Unlike prior models with recurrent or stacked temporal encoders, SOTFormer achieves stable identity propagation through a ground-truth-primed memory and a burn-in anchor loss that explicitly stabilizes initialization. A single lightweight temporal-attention layer refines embeddings across frames, enabling real-time inference with fixed GPU memory. On the Mini-LaSOT (20%) benchmark, SOTFormer attains 76.3 AUC and 53.7 FPS (AMP, 4.3 GB VRAM), outperforming transformer baselines such as TrackFormer and MOTRv2 under fast motion, scale change, and occlusion.
title SOTFormer: A Minimal Transformer for Unified Object Tracking and Trajectory Prediction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.11824