Unified Multimodal Visual Tracking with Dual Mixture-of-Experts
Fuente:
arXiv
Saved in:
| Main Authors: | Hong, Lingyi, Li, Jinglun, Zhou, Xinyu, Jiang, Kaixun, Guo, Pinxue, Chen, Zhaoyu, Li, Runze, Sheng, Xingdong, Zhang, Wenqiang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DeTrack: In-model Latent Denoising Learning for Visual Object Tracking
by: Zhou, Xinyu, et al.
Published: (2025)
by: Zhou, Xinyu, et al.
Published: (2025)
OneTracker: Unifying Visual Object Tracking with Foundation Models and Efficient Tuning
by: Hong, Lingyi, et al.
Published: (2024)
by: Hong, Lingyi, et al.
Published: (2024)
General Compression Framework for Efficient Transformer Object Tracking
by: Hong, Lingyi, et al.
Published: (2024)
by: Hong, Lingyi, et al.
Published: (2024)
TagOOD: A Novel Approach to Out-of-Distribution Detection via Vision-Language Representations and Class Center Learning
by: Li, Jinglun, et al.
Published: (2024)
by: Li, Jinglun, et al.
Published: (2024)
X-Prompt: Multi-modal Visual Prompt for Video Object Segmentation
by: Guo, Pinxue, et al.
Published: (2024)
by: Guo, Pinxue, et al.
Published: (2024)
Dynamic Semantic-Aware Correlation Modeling for UAV Tracking
by: Zhou, Xinyu, et al.
Published: (2025)
by: Zhou, Xinyu, et al.
Published: (2025)
Reading Relevant Feature from Global Representation Memory for Visual Object Tracking
by: Zhou, Xinyu, et al.
Published: (2024)
by: Zhou, Xinyu, et al.
Published: (2024)
Improving Adversarial Transferability with Neighbourhood Gradient Information
by: Guo, Haijing, et al.
Published: (2024)
by: Guo, Haijing, et al.
Published: (2024)
Seeing is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding
by: Guo, Pinxue, et al.
Published: (2025)
by: Guo, Pinxue, et al.
Published: (2025)
VideoPure: Diffusion-based Adversarial Purification for Video Recognition
by: Jiang, Kaixun, et al.
Published: (2025)
by: Jiang, Kaixun, et al.
Published: (2025)
ClickVOS: Click Video Object Segmentation
by: Guo, Pinxue, et al.
Published: (2024)
by: Guo, Pinxue, et al.
Published: (2024)
Enhancing Diffusion-based Unrestricted Adversarial Attacks via Adversary Preferences Alignment
by: Jiang, Kaixun, et al.
Published: (2025)
by: Jiang, Kaixun, et al.
Published: (2025)
Boosting the Transferability of Adversarial Attacks with Global Momentum Initialization
by: Wang, Jiafeng, et al.
Published: (2022)
by: Wang, Jiafeng, et al.
Published: (2022)
Hierarchical Visual Categories Modeling: A Joint Representation Learning and Density Estimation Framework for Out-of-Distribution Detection
by: Li, Jinglun, et al.
Published: (2024)
by: Li, Jinglun, et al.
Published: (2024)
OneVOS: Unifying Video Object Segmentation with All-in-One Transformer Framework
by: Li, Wanyun, et al.
Published: (2024)
by: Li, Wanyun, et al.
Published: (2024)
LVOS: A Benchmark for Large-scale Long-term Video Object Segmentation
by: Hong, Lingyi, et al.
Published: (2024)
by: Hong, Lingyi, et al.
Published: (2024)
Synthesizing Near-Boundary OOD Samples for Out-of-Distribution Detection
by: Li, Jinglun, et al.
Published: (2025)
by: Li, Jinglun, et al.
Published: (2025)
VLA-Hijack: A Transferable Patch Attack against Vision-Language-Action Models via Visual Proprioception Hijacking
by: Fu, Jiyuan, et al.
Published: (2026)
by: Fu, Jiyuan, et al.
Published: (2026)
PG-Attack: A Precision-Guided Adversarial Attack Framework Against Vision Foundation Models for Autonomous Driving
by: Fu, Jiyuan, et al.
Published: (2024)
by: Fu, Jiyuan, et al.
Published: (2024)
GenAgent: Scaling Text-to-Image Generation via Agentic Multimodal Reasoning
by: Jiang, Kaixun, et al.
Published: (2026)
by: Jiang, Kaixun, et al.
Published: (2026)
Improving Adversarial Transferability of Vision-Language Pre-training Models through Collaborative Multimodal Interaction
by: Fu, Jiyuan, et al.
Published: (2024)
by: Fu, Jiyuan, et al.
Published: (2024)
Boosting Adversarial Transferability with Spatial Adversarial Alignment
by: Chen, Zhaoyu, et al.
Published: (2025)
by: Chen, Zhaoyu, et al.
Published: (2025)
Exploring the Adversarial Robustness of Face Forgery Detection with Decision-based Black-box Attacks
by: Chen, Zhaoyu, et al.
Published: (2023)
by: Chen, Zhaoyu, et al.
Published: (2023)
OpenVIS: Open-vocabulary Video Instance Segmentation
by: Guo, Pinxue, et al.
Published: (2023)
by: Guo, Pinxue, et al.
Published: (2023)
Gamma: Toward Generic Image Assessment with Mixture of Assessment Experts
by: Zhou, Hantao, et al.
Published: (2025)
by: Zhou, Hantao, et al.
Published: (2025)
RSAgent: Learning to Reason and Act for Text-Guided Segmentation via Multi-Turn Tool Invocations
by: He, Xingqi, et al.
Published: (2025)
by: He, Xingqi, et al.
Published: (2025)
Improving Multimodal Sentiment Analysis via Modality Optimization and Dynamic Primary Modality Selection
by: Yang, Dingkang, et al.
Published: (2025)
by: Yang, Dingkang, et al.
Published: (2025)
Delving into Decision-based Black-box Attacks on Semantic Segmentation
by: Chen, Zhaoyu, et al.
Published: (2024)
by: Chen, Zhaoyu, et al.
Published: (2024)
Sparse-Dense Mixture of Experts Adapter for Multi-Modal Tracking
by: Zhu, Yabin, et al.
Published: (2026)
by: Zhu, Yabin, et al.
Published: (2026)
Aria: An Open Multimodal Native Mixture-of-Experts Model
by: Li, Dongxu, et al.
Published: (2024)
by: Li, Dongxu, et al.
Published: (2024)
MoVA: Adapting Mixture of Vision Experts to Multimodal Context
by: Zong, Zhuofan, et al.
Published: (2024)
by: Zong, Zhuofan, et al.
Published: (2024)
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
CL-MoE: Enhancing Multimodal Large Language Model with Dual Momentum Mixture-of-Experts for Continual Visual Question Answering
by: Huai, Tianyu, et al.
Published: (2025)
by: Huai, Tianyu, et al.
Published: (2025)
Towards General Multimodal Visual Tracking
by: Lu, Andong, et al.
Published: (2025)
by: Lu, Andong, et al.
Published: (2025)
Towards Unified Multimodal Editing with Enhanced Knowledge Collaboration
by: Pan, Kaihang, et al.
Published: (2024)
by: Pan, Kaihang, et al.
Published: (2024)
TAG-MoE: Task-Aware Gating for Unified Generative Mixture-of-Experts
by: Xu, Yu, et al.
Published: (2026)
by: Xu, Yu, et al.
Published: (2026)
UAUTrack: Towards Unified Multimodal Anti-UAV Visual Tracking
by: Ren, Qionglin, et al.
Published: (2025)
by: Ren, Qionglin, et al.
Published: (2025)
MSVCOD:A Large-Scale Multi-Scene Dataset for Video Camouflage Object Detection
by: Gao, Shuyong, et al.
Published: (2025)
by: Gao, Shuyong, et al.
Published: (2025)
SPMTrack: Spatio-Temporal Parameter-Efficient Fine-Tuning with Mixture of Experts for Scalable Visual Tracking
by: Cai, Wenrui, et al.
Published: (2025)
by: Cai, Wenrui, et al.
Published: (2025)
Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts
by: Xu, Haolei, et al.
Published: (2026)
by: Xu, Haolei, et al.
Published: (2026)
Similar Items
-
DeTrack: In-model Latent Denoising Learning for Visual Object Tracking
by: Zhou, Xinyu, et al.
Published: (2025) -
OneTracker: Unifying Visual Object Tracking with Foundation Models and Efficient Tuning
by: Hong, Lingyi, et al.
Published: (2024) -
General Compression Framework for Efficient Transformer Object Tracking
by: Hong, Lingyi, et al.
Published: (2024) -
TagOOD: A Novel Approach to Out-of-Distribution Detection via Vision-Language Representations and Class Center Learning
by: Li, Jinglun, et al.
Published: (2024) -
X-Prompt: Multi-modal Visual Prompt for Video Object Segmentation
by: Guo, Pinxue, et al.
Published: (2024)