TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yang, Fan, Zheng, Shurong, Zhao, Hongyin, Zhan, Yufei, Li, Xin, Zhu, Yousong, Tang, Chaoyang Zhao Ming, Wang, Jinqiao
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910031055683584
author Yang, Fan
Zheng, Shurong
Zhao, Hongyin
Zhan, Yufei
Li, Xin
Zhu, Yousong
Tang, Chaoyang Zhao Ming
Wang, Jinqiao
author_facet Yang, Fan
Zheng, Shurong
Zhao, Hongyin
Zhan, Yufei
Li, Xin
Zhu, Yousong
Tang, Chaoyang Zhao Ming
Wang, Jinqiao
contents Recent Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities in image understanding and natural language generation. However, current approaches focus predominantly on global image understanding, struggling to simulate human visual attention trajectories and explain associations between descriptions and specific regions. We propose TraceVision, a unified vision-language model integrating trajectory-aware spatial understanding in an end-to-end framework. TraceVision employs a Trajectory-aware Visual Perception (TVP) module for bidirectional fusion of visual features and trajectory information. We design geometric simplification to extract semantic keypoints from raw trajectories and propose a three-stage training pipeline where trajectories guide description generation and region localization. We extend TraceVision to trajectory-guided segmentation and video scene understanding, enabling cross-frame tracking and temporal attention analysis. We construct the Reasoning-based Interactive Localized Narratives (RILN) dataset to enhance logical reasoning and interpretability. Extensive experiments on trajectory-guided captioning, text-guided trajectory prediction, understanding, and segmentation demonstrate that TraceVision achieves state-of-the-art performance, establishing a foundation for intuitive spatial interaction and interpretable visual understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2602_19768
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding
Yang, Fan
Zheng, Shurong
Zhao, Hongyin
Zhan, Yufei
Li, Xin
Zhu, Yousong
Tang, Chaoyang Zhao Ming
Wang, Jinqiao
Computer Vision and Pattern Recognition
Recent Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities in image understanding and natural language generation. However, current approaches focus predominantly on global image understanding, struggling to simulate human visual attention trajectories and explain associations between descriptions and specific regions. We propose TraceVision, a unified vision-language model integrating trajectory-aware spatial understanding in an end-to-end framework. TraceVision employs a Trajectory-aware Visual Perception (TVP) module for bidirectional fusion of visual features and trajectory information. We design geometric simplification to extract semantic keypoints from raw trajectories and propose a three-stage training pipeline where trajectories guide description generation and region localization. We extend TraceVision to trajectory-guided segmentation and video scene understanding, enabling cross-frame tracking and temporal attention analysis. We construct the Reasoning-based Interactive Localized Narratives (RILN) dataset to enhance logical reasoning and interpretability. Extensive experiments on trajectory-guided captioning, text-guided trajectory prediction, understanding, and segmentation demonstrate that TraceVision achieves state-of-the-art performance, establishing a foundation for intuitive spatial interaction and interpretable visual understanding.
title TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.19768