CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Fiastre, Gabriel, Yang, Antoine, Schmid, Cordelia
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914622441783296
author Fiastre, Gabriel
Yang, Antoine
Schmid, Cordelia
author_facet Fiastre, Gabriel
Yang, Antoine
Schmid, Cordelia
contents Dense Video Object Captioning (DVOC) is the task of jointly detecting, tracking, and captioning object trajectories in a video, requiring the ability to understand spatio-temporal details and describe them in natural language. Due to the complexity of the task and the high cost associated with manual annotation, previous approaches resort to training strategies with limited data, potentially leading to suboptimal performance. To circumvent this issue, we propose to generate captions about spatio-temporally localized entities leveraging a state-of-the-art VLM, and extend the LVIS and LV-VIS datasets with our synthetic captions (LVISCap and LV-VISCap). Moreover, we introduce an end-to-end model, CaptionFormer, capable of jointly detecting, segmenting, tracking and captioning object trajectories. CaptionFormer achieves state-of-the-art DVOC results on three existing benchmarks, VidSTG, VLN and BenSMOT. The datasets and code are available at https://www.gabriel.fiastre.fr/captionformer/.
format Preprint
id arxiv_https___arxiv_org_abs_2510_14904
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
Fiastre, Gabriel
Yang, Antoine
Schmid, Cordelia
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Dense Video Object Captioning (DVOC) is the task of jointly detecting, tracking, and captioning object trajectories in a video, requiring the ability to understand spatio-temporal details and describe them in natural language. Due to the complexity of the task and the high cost associated with manual annotation, previous approaches resort to training strategies with limited data, potentially leading to suboptimal performance. To circumvent this issue, we propose to generate captions about spatio-temporally localized entities leveraging a state-of-the-art VLM, and extend the LVIS and LV-VIS datasets with our synthetic captions (LVISCap and LV-VISCap). Moreover, we introduce an end-to-end model, CaptionFormer, capable of jointly detecting, segmenting, tracking and captioning object trajectories. CaptionFormer achieves state-of-the-art DVOC results on three existing benchmarks, VidSTG, VLN and BenSMOT. The datasets and code are available at https://www.gabriel.fiastre.fr/captionformer/.
title CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.14904