TrackVLA: Embodied Visual Tracking in the Wild

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Shaoan, Zhang, Jiazhao, Li, Minghan, Liu, Jiahang, Li, Anqi, Wu, Kui, Zhong, Fangwei, Yu, Junzhi, Zhang, Zhizheng, Wang, He
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910973647912960
author Wang, Shaoan
Zhang, Jiazhao
Li, Minghan
Liu, Jiahang
Li, Anqi
Wu, Kui
Zhong, Fangwei
Yu, Junzhi
Zhang, Zhizheng
Wang, He
author_facet Wang, Shaoan
Zhang, Jiazhao
Li, Minghan
Liu, Jiahang
Li, Anqi
Wu, Kui
Zhong, Fangwei
Yu, Junzhi
Zhang, Zhizheng
Wang, He
contents Embodied visual tracking is a fundamental skill in Embodied AI, enabling an agent to follow a specific target in dynamic environments using only egocentric vision. This task is inherently challenging as it requires both accurate target recognition and effective trajectory planning under conditions of severe occlusion and high scene dynamics. Existing approaches typically address this challenge through a modular separation of recognition and planning. In this work, we propose TrackVLA, a Vision-Language-Action (VLA) model that learns the synergy between object recognition and trajectory planning. Leveraging a shared LLM backbone, we employ a language modeling head for recognition and an anchor-based diffusion model for trajectory planning. To train TrackVLA, we construct an Embodied Visual Tracking Benchmark (EVT-Bench) and collect diverse difficulty levels of recognition samples, resulting in a dataset of 1.7 million samples. Through extensive experiments in both synthetic and real-world environments, TrackVLA demonstrates SOTA performance and strong generalizability. It significantly outperforms existing methods on public benchmarks in a zero-shot manner while remaining robust to high dynamics and occlusion in real-world scenarios at 10 FPS inference speed. Our project page is: https://pku-epic.github.io/TrackVLA-web.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23189
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TrackVLA: Embodied Visual Tracking in the Wild
Wang, Shaoan
Zhang, Jiazhao
Li, Minghan
Liu, Jiahang
Li, Anqi
Wu, Kui
Zhong, Fangwei
Yu, Junzhi
Zhang, Zhizheng
Wang, He
Robotics
Computer Vision and Pattern Recognition
Embodied visual tracking is a fundamental skill in Embodied AI, enabling an agent to follow a specific target in dynamic environments using only egocentric vision. This task is inherently challenging as it requires both accurate target recognition and effective trajectory planning under conditions of severe occlusion and high scene dynamics. Existing approaches typically address this challenge through a modular separation of recognition and planning. In this work, we propose TrackVLA, a Vision-Language-Action (VLA) model that learns the synergy between object recognition and trajectory planning. Leveraging a shared LLM backbone, we employ a language modeling head for recognition and an anchor-based diffusion model for trajectory planning. To train TrackVLA, we construct an Embodied Visual Tracking Benchmark (EVT-Bench) and collect diverse difficulty levels of recognition samples, resulting in a dataset of 1.7 million samples. Through extensive experiments in both synthetic and real-world environments, TrackVLA demonstrates SOTA performance and strong generalizability. It significantly outperforms existing methods on public benchmarks in a zero-shot manner while remaining robust to high dynamics and occlusion in real-world scenarios at 10 FPS inference speed. Our project page is: https://pku-epic.github.io/TrackVLA-web.
title TrackVLA: Embodied Visual Tracking in the Wild
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.23189