STORM: End-to-End Referring Multi-Object Tracking in Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Zijia, Yi, Jingru, Wang, Jue, Chen, Yuxiao, Chen, Junwen, Li, Xinyu, Modolo, Davide
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911586681094144
author Lu, Zijia
Yi, Jingru
Wang, Jue
Chen, Yuxiao
Chen, Junwen
Li, Xinyu
Modolo, Davide
author_facet Lu, Zijia
Yi, Jingru
Wang, Jue
Chen, Yuxiao
Chen, Junwen
Li, Xinyu
Modolo, Davide
contents Referring multi-object tracking (RMOT) is a task of associating all the objects in a video that semantically match with given textual queries or referring expressions. Existing RMOT approaches decompose object grounding and tracking into separated modules and exhibit limited performance due to the scarcity of training videos, ambiguous annotations, and restricted domains. In this work, we introduce STORM, an end-to-end MLLM that jointly performs grounding and tracking within a unified framework, eliminating external detectors and enabling coherent reasoning over appearance, motion, and language. To improve data efficiency, we propose a task-composition learning (TCL) strategy that decomposes RMOT into image grounding and object tracking, allowing STORM to leverage data-rich sub-tasks and learn structured spatial--temporal reasoning. We further construct STORM-Bench, a new RMOT dataset with accurate trajectories and diverse, unambiguous referring expressions generated through a bottom-up annotation pipeline. Extensive experiments show that STORM achieves state-of-the-art performance on image grounding, single-object tracking, and RMOT benchmarks, demonstrating strong generalization and robust spatial--temporal grounding in complex real-world scenarios. STORM-Bench is released at https://github.com/amazon-science/storm-referring-multi-object-grounding.
format Preprint
id arxiv_https___arxiv_org_abs_2604_10527
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle STORM: End-to-End Referring Multi-Object Tracking in Videos
Lu, Zijia
Yi, Jingru
Wang, Jue
Chen, Yuxiao
Chen, Junwen
Li, Xinyu
Modolo, Davide
Computer Vision and Pattern Recognition
Artificial Intelligence
Referring multi-object tracking (RMOT) is a task of associating all the objects in a video that semantically match with given textual queries or referring expressions. Existing RMOT approaches decompose object grounding and tracking into separated modules and exhibit limited performance due to the scarcity of training videos, ambiguous annotations, and restricted domains. In this work, we introduce STORM, an end-to-end MLLM that jointly performs grounding and tracking within a unified framework, eliminating external detectors and enabling coherent reasoning over appearance, motion, and language. To improve data efficiency, we propose a task-composition learning (TCL) strategy that decomposes RMOT into image grounding and object tracking, allowing STORM to leverage data-rich sub-tasks and learn structured spatial--temporal reasoning. We further construct STORM-Bench, a new RMOT dataset with accurate trajectories and diverse, unambiguous referring expressions generated through a bottom-up annotation pipeline. Extensive experiments show that STORM achieves state-of-the-art performance on image grounding, single-object tracking, and RMOT benchmarks, demonstrating strong generalization and robust spatial--temporal grounding in complex real-world scenarios. STORM-Bench is released at https://github.com/amazon-science/storm-referring-multi-object-grounding.
title STORM: End-to-End Referring Multi-Object Tracking in Videos
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2604.10527