GLAD: Generative Language-Assisted Visual Tracking for Low-Semantic Templates

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Luo, Xingyu, Cai, Yidong, Liu, Jie, Tang, Jie, Wu, Gangshan, Wang, Limin
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914298066894848
author Luo, Xingyu
Cai, Yidong
Liu, Jie
Tang, Jie
Wu, Gangshan
Wang, Limin
author_facet Luo, Xingyu
Cai, Yidong
Liu, Jie
Tang, Jie
Wu, Gangshan
Wang, Limin
contents Vision-language tracking has gained increasing attention in many scenarios. This task simultaneously deals with visual and linguistic information to localize objects in videos. Despite its growing utility, the development of vision-language tracking methods remains in its early stage. Current vision-language trackers usually employ Transformer architectures for interactive integration of template, search, and text features. However, persistent challenges about low-semantic images including prevalent image blurriness, low resolution and so on, may compromise model performance through degraded cross-modal understanding. To solve this problem, language assistance is usually used to deal with the obstacles posed by low-semantic images. However, due to the existing gap between current textual and visual features, direct concatenation and fusion of these features may have limited effectiveness. To address these challenges, we introduce a pioneering Generative Language-AssisteD tracking model, GLAD, which utilizes diffusion models for the generative multi-modal fusion of text description and template image to bolster compatibility between language and image and enhance template image semantic information. Our approach demonstrates notable improvements over the existing fusion paradigms. Blurry and semantically ambiguous template images can be restored to improve multi-modal features in the generative fusion paradigm. Experiments show that our method establishes a new state-of-the-art on multiple benchmarks and achieves an impressive inference speed. The code and models will be released at: https://github.com/Confetti-lxy/GLAD
format Preprint
id arxiv_https___arxiv_org_abs_2602_00570
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GLAD: Generative Language-Assisted Visual Tracking for Low-Semantic Templates
Luo, Xingyu
Cai, Yidong
Liu, Jie
Tang, Jie
Wu, Gangshan
Wang, Limin
Computer Vision and Pattern Recognition
Vision-language tracking has gained increasing attention in many scenarios. This task simultaneously deals with visual and linguistic information to localize objects in videos. Despite its growing utility, the development of vision-language tracking methods remains in its early stage. Current vision-language trackers usually employ Transformer architectures for interactive integration of template, search, and text features. However, persistent challenges about low-semantic images including prevalent image blurriness, low resolution and so on, may compromise model performance through degraded cross-modal understanding. To solve this problem, language assistance is usually used to deal with the obstacles posed by low-semantic images. However, due to the existing gap between current textual and visual features, direct concatenation and fusion of these features may have limited effectiveness. To address these challenges, we introduce a pioneering Generative Language-AssisteD tracking model, GLAD, which utilizes diffusion models for the generative multi-modal fusion of text description and template image to bolster compatibility between language and image and enhance template image semantic information. Our approach demonstrates notable improvements over the existing fusion paradigms. Blurry and semantically ambiguous template images can be restored to improve multi-modal features in the generative fusion paradigm. Experiments show that our method establishes a new state-of-the-art on multiple benchmarks and achieves an impressive inference speed. The code and models will be released at: https://github.com/Confetti-lxy/GLAD
title GLAD: Generative Language-Assisted Visual Tracking for Low-Semantic Templates
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.00570