Context-Aware Integration of Language and Visual References for Natural Language Tracking

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shao, Yanyan, He, Shuting, Ye, Qi, Feng, Yuchao, Luo, Wenhan, Chen, Jiming
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911818992058368
author Shao, Yanyan
He, Shuting
Ye, Qi
Feng, Yuchao
Luo, Wenhan
Chen, Jiming
author_facet Shao, Yanyan
He, Shuting
Ye, Qi
Feng, Yuchao
Luo, Wenhan
Chen, Jiming
contents Tracking by natural language specification (TNL) aims to consistently localize a target in a video sequence given a linguistic description in the initial frame. Existing methodologies perform language-based and template-based matching for target reasoning separately and merge the matching results from two sources, which suffer from tracking drift when language and visual templates miss-align with the dynamic target state and ambiguity in the later merging stage. To tackle the issues, we propose a joint multi-modal tracking framework with 1) a prompt modulation module to leverage the complementarity between temporal visual templates and language expressions, enabling precise and context-aware appearance and linguistic cues, and 2) a unified target decoding module to integrate the multi-modal reference cues and executes the integrated queries on the search image to predict the target location in an end-to-end manner directly. This design ensures spatio-temporal consistency by leveraging historical visual information and introduces an integrated solution, generating predictions in a single step. Extensive experiments conducted on TNL2K, OTB-Lang, LaSOT, and RefCOCOg validate the efficacy of our proposed approach. The results demonstrate competitive performance against state-of-the-art methods for both tracking and grounding.
format Preprint
id arxiv_https___arxiv_org_abs_2403_19975
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Context-Aware Integration of Language and Visual References for Natural Language Tracking
Shao, Yanyan
He, Shuting
Ye, Qi
Feng, Yuchao
Luo, Wenhan
Chen, Jiming
Computer Vision and Pattern Recognition
Tracking by natural language specification (TNL) aims to consistently localize a target in a video sequence given a linguistic description in the initial frame. Existing methodologies perform language-based and template-based matching for target reasoning separately and merge the matching results from two sources, which suffer from tracking drift when language and visual templates miss-align with the dynamic target state and ambiguity in the later merging stage. To tackle the issues, we propose a joint multi-modal tracking framework with 1) a prompt modulation module to leverage the complementarity between temporal visual templates and language expressions, enabling precise and context-aware appearance and linguistic cues, and 2) a unified target decoding module to integrate the multi-modal reference cues and executes the integrated queries on the search image to predict the target location in an end-to-end manner directly. This design ensures spatio-temporal consistency by leveraging historical visual information and introduces an integrated solution, generating predictions in a single step. Extensive experiments conducted on TNL2K, OTB-Lang, LaSOT, and RefCOCOg validate the efficacy of our proposed approach. The results demonstrate competitive performance against state-of-the-art methods for both tracking and grounding.
title Context-Aware Integration of Language and Visual References for Natural Language Tracking
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.19975