Unified Multimodal Visual Tracking with Dual Mixture-of-Experts

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Hong, Lingyi, Li, Jinglun, Zhou, Xinyu, Jiang, Kaixun, Guo, Pinxue, Chen, Zhaoyu, Li, Runze, Sheng, Xingdong, Zhang, Wenqiang
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918484797030400
author Hong, Lingyi
Li, Jinglun
Zhou, Xinyu
Jiang, Kaixun
Guo, Pinxue
Chen, Zhaoyu
Li, Runze
Sheng, Xingdong
Zhang, Wenqiang
author_facet Hong, Lingyi
Li, Jinglun
Zhou, Xinyu
Jiang, Kaixun
Guo, Pinxue
Chen, Zhaoyu
Li, Runze
Sheng, Xingdong
Zhang, Wenqiang
contents Multimodal visual object tracking can be divided into to several kinds of tasks (e.g. RGB and RGB+X tracking), based on the input modality. Existing methods often train separate models for each modality or rely on pretrained models to adapt to new modalities, which limits efficiency, scalability, and usability. Thus, we introduce OneTrackerV2, a unified multi-modal tracking framework that enables end-to-end training for any modality. We propose Meta Merger to embed multi-modal information into a unified space, allowing flexible modality fusion and robustness. We further introduce Dual Mixture-of-Experts (DMoE): T-MoE models spatio-temporal relations for tracking, while M-MoE embeds multi-modal knowledge, disentangling cross-modal dependencies and reducing feature conflicts. With a shared architecture, unified parameters, and a single end-to-end training, OneTrackerV2 achieves state-of-the-art performance across five RGB and RGB+X tracking tasks and 12 benchmarks, while maintaining high inference efficiency. Notably, even after model compression, OneTrackerV2 retains strong performance. Moreover, OneTrackerV2 demonstrates remarkable robustness under modality-missing scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2605_03716
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Unified Multimodal Visual Tracking with Dual Mixture-of-Experts
Hong, Lingyi
Li, Jinglun
Zhou, Xinyu
Jiang, Kaixun
Guo, Pinxue
Chen, Zhaoyu
Li, Runze
Sheng, Xingdong
Zhang, Wenqiang
Computer Vision and Pattern Recognition
Multimodal visual object tracking can be divided into to several kinds of tasks (e.g. RGB and RGB+X tracking), based on the input modality. Existing methods often train separate models for each modality or rely on pretrained models to adapt to new modalities, which limits efficiency, scalability, and usability. Thus, we introduce OneTrackerV2, a unified multi-modal tracking framework that enables end-to-end training for any modality. We propose Meta Merger to embed multi-modal information into a unified space, allowing flexible modality fusion and robustness. We further introduce Dual Mixture-of-Experts (DMoE): T-MoE models spatio-temporal relations for tracking, while M-MoE embeds multi-modal knowledge, disentangling cross-modal dependencies and reducing feature conflicts. With a shared architecture, unified parameters, and a single end-to-end training, OneTrackerV2 achieves state-of-the-art performance across five RGB and RGB+X tracking tasks and 12 benchmarks, while maintaining high inference efficiency. Notably, even after model compression, OneTrackerV2 retains strong performance. Moreover, OneTrackerV2 demonstrates remarkable robustness under modality-missing scenarios.
title Unified Multimodal Visual Tracking with Dual Mixture-of-Experts
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.03716