Cognitive Disentanglement for Referring Multi-Object Tracking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Shaofeng, Guan, Runwei, Lian, Wangwang, Liu, Daizong, Sun, Xiaolou, Wu, Dongming, Yue, Yutao, Ding, Weiping, Xiong, Hui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915306970021888
author Liang, Shaofeng
Guan, Runwei
Lian, Wangwang
Liu, Daizong
Sun, Xiaolou
Wu, Dongming
Yue, Yutao
Ding, Weiping
Xiong, Hui
author_facet Liang, Shaofeng
Guan, Runwei
Lian, Wangwang
Liu, Daizong
Sun, Xiaolou
Wu, Dongming
Yue, Yutao
Ding, Weiping
Xiong, Hui
contents As a significant application of multi-source information fusion in intelligent transportation perception systems, Referring Multi-Object Tracking (RMOT) involves localizing and tracking specific objects in video sequences based on language references. However, existing RMOT approaches often treat language descriptions as holistic embeddings and struggle to effectively integrate the rich semantic information contained in language expressions with visual features. This limitation is especially apparent in complex scenes requiring comprehensive understanding of both static object attributes and spatial motion information. In this paper, we propose a Cognitive Disentanglement for Referring Multi-Object Tracking (CDRMT) framework that addresses these challenges. It adapts the "what" and "where" pathways from the human visual processing system to RMOT tasks. Specifically, our framework first establishes cross-modal connections while preserving modality-specific characteristics. It then disentangles language descriptions and hierarchically injects them into object queries, refining object understanding from coarse to fine-grained semantic levels. Finally, we reconstruct language representations based on visual features, ensuring that tracked objects faithfully reflect the referring expression. Extensive experiments on different benchmark datasets demonstrate that CDRMT achieves substantial improvements over state-of-the-art methods, with average gains of 6.0% in HOTA score on Refer-KITTI and 3.2% on Refer-KITTI-V2. Our approach advances the state-of-the-art in RMOT while simultaneously providing new insights into multi-source information fusion.
format Preprint
id arxiv_https___arxiv_org_abs_2503_11496
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Cognitive Disentanglement for Referring Multi-Object Tracking
Liang, Shaofeng
Guan, Runwei
Lian, Wangwang
Liu, Daizong
Sun, Xiaolou
Wu, Dongming
Yue, Yutao
Ding, Weiping
Xiong, Hui
Computer Vision and Pattern Recognition
As a significant application of multi-source information fusion in intelligent transportation perception systems, Referring Multi-Object Tracking (RMOT) involves localizing and tracking specific objects in video sequences based on language references. However, existing RMOT approaches often treat language descriptions as holistic embeddings and struggle to effectively integrate the rich semantic information contained in language expressions with visual features. This limitation is especially apparent in complex scenes requiring comprehensive understanding of both static object attributes and spatial motion information. In this paper, we propose a Cognitive Disentanglement for Referring Multi-Object Tracking (CDRMT) framework that addresses these challenges. It adapts the "what" and "where" pathways from the human visual processing system to RMOT tasks. Specifically, our framework first establishes cross-modal connections while preserving modality-specific characteristics. It then disentangles language descriptions and hierarchically injects them into object queries, refining object understanding from coarse to fine-grained semantic levels. Finally, we reconstruct language representations based on visual features, ensuring that tracked objects faithfully reflect the referring expression. Extensive experiments on different benchmark datasets demonstrate that CDRMT achieves substantial improvements over state-of-the-art methods, with average gains of 6.0% in HOTA score on Refer-KITTI and 3.2% on Refer-KITTI-V2. Our approach advances the state-of-the-art in RMOT while simultaneously providing new insights into multi-source information fusion.
title Cognitive Disentanglement for Referring Multi-Object Tracking
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.11496