Temporal Object Captioning for Street Scene Videos from LiDAR Tracks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gopinathan, Vignesh, Zimmermann, Urs, Arnold, Michael, Rottmann, Matthias
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913852496543744
author Gopinathan, Vignesh
Zimmermann, Urs
Arnold, Michael
Rottmann, Matthias
author_facet Gopinathan, Vignesh
Zimmermann, Urs
Arnold, Michael
Rottmann, Matthias
contents Video captioning models have seen notable advancements in recent years, especially with regard to their ability to capture temporal information. While many research efforts have focused on architectural advancements, such as temporal attention mechanisms, there remains a notable gap in understanding how models capture and utilize temporal semantics for effective temporal feature extraction, especially in the context of Advanced Driver Assistance Systems. We propose an automated LiDAR-based captioning procedure that focuses on the temporal dynamics of traffic participants. Our approach uses a rule-based system to extract essential details such as lane position and relative motion from object tracks, followed by a template-based caption generation. Our findings show that training SwinBERT, a video captioning model, using only front camera images and supervised with our template-based captions, specifically designed to encapsulate fine-grained temporal behavior, leads to improved temporal understanding consistently across three datasets. In conclusion, our results clearly demonstrate that integrating LiDAR-based caption supervision significantly enhances temporal understanding, effectively addressing and reducing the inherent visual/static biases prevalent in current state-of-the-art model architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16594
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Temporal Object Captioning for Street Scene Videos from LiDAR Tracks
Gopinathan, Vignesh
Zimmermann, Urs
Arnold, Michael
Rottmann, Matthias
Computer Vision and Pattern Recognition
Machine Learning
Video captioning models have seen notable advancements in recent years, especially with regard to their ability to capture temporal information. While many research efforts have focused on architectural advancements, such as temporal attention mechanisms, there remains a notable gap in understanding how models capture and utilize temporal semantics for effective temporal feature extraction, especially in the context of Advanced Driver Assistance Systems. We propose an automated LiDAR-based captioning procedure that focuses on the temporal dynamics of traffic participants. Our approach uses a rule-based system to extract essential details such as lane position and relative motion from object tracks, followed by a template-based caption generation. Our findings show that training SwinBERT, a video captioning model, using only front camera images and supervised with our template-based captions, specifically designed to encapsulate fine-grained temporal behavior, leads to improved temporal understanding consistently across three datasets. In conclusion, our results clearly demonstrate that integrating LiDAR-based caption supervision significantly enhances temporal understanding, effectively addressing and reducing the inherent visual/static biases prevalent in current state-of-the-art model architectures.
title Temporal Object Captioning for Street Scene Videos from LiDAR Tracks
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2505.16594