SEATrack: Simple, Efficient, and Adaptive Multimodal Tracker
Fuente:
arXiv
Saved in:
| Main Authors: | Su, Junbin, Xue, Ziteng, Zhang, Shihui, Chen, Kun, Hu, Weiming, Zhang, Zhipeng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DARTer: Dynamic Adaptive Representation Tracker for Nighttime UAV Tracking
by: Li, Xuzhao, et al.
Published: (2025)
by: Li, Xuzhao, et al.
Published: (2025)
PhyTracker: An Online Tracker for Phytoplankton
by: Yu, Yang, et al.
Published: (2024)
by: Yu, Yang, et al.
Published: (2024)
Geometry OR Tracker: Universal Geometric Operating Room Tracking
by: Shao, Yihua, et al.
Published: (2026)
by: Shao, Yihua, et al.
Published: (2026)
A Simple Aerial Detection Baseline of Multimodal Language Models
by: Li, Qingyun, et al.
Published: (2025)
by: Li, Qingyun, et al.
Published: (2025)
A Simple Background Augmentation Method for Object Detection with Diffusion Model
by: Li, Yuhang, et al.
Published: (2024)
by: Li, Yuhang, et al.
Published: (2024)
AdaTok: Adaptive Token Compression with Object-Aware Representations for Efficient Multimodal LLMs
by: Zhang, Xinliang, et al.
Published: (2025)
by: Zhang, Xinliang, et al.
Published: (2025)
Ada-Tracker: Soft Tissue Tracking via Inter-Frame and Adaptive-Template Matching
by: Guo, Jiaxin, et al.
Published: (2024)
by: Guo, Jiaxin, et al.
Published: (2024)
ResAdapt: Adaptive Resolution for Efficient Multimodal Reasoning
by: Liao, Huanxuan, et al.
Published: (2026)
by: Liao, Huanxuan, et al.
Published: (2026)
SWIFT: Prompt-Adaptive Memory for Efficient Interactive Long Video Generation
by: Tan, Shanwen, et al.
Published: (2026)
by: Tan, Shanwen, et al.
Published: (2026)
Multimodal Prompt Decoupling Attack on the Safety Filters in Text-to-Image Models
by: Peng, Xingkai, et al.
Published: (2025)
by: Peng, Xingkai, et al.
Published: (2025)
Sparse Shortcuts: Facilitating Efficient Fusion in Multimodal Large Language Models
by: Zhang, Jingrui, et al.
Published: (2026)
by: Zhang, Jingrui, et al.
Published: (2026)
Complementary Pseudo Multimodal Feature for Point Cloud Anomaly Detection
by: Cao, Yunkang, et al.
Published: (2023)
by: Cao, Yunkang, et al.
Published: (2023)
KAN-RCBEVDepth: A multi-modal fusion algorithm in object detection for autonomous driving
by: Lai, Zhihao, et al.
Published: (2024)
by: Lai, Zhihao, et al.
Published: (2024)
Adaptive Clinical-Aware Latent Diffusion for Multimodal Brain Image Generation and Missing Modality Imputation
by: Zhou, Rong, et al.
Published: (2026)
by: Zhou, Rong, et al.
Published: (2026)
AVG-LLaVA: An Efficient Large Multimodal Model with Adaptive Visual Granularity
by: Lan, Zhibin, et al.
Published: (2024)
by: Lan, Zhibin, et al.
Published: (2024)
MM-DeepResearch: A Simple and Effective Multimodal Agentic Search Baseline
by: Yao, Huanjin, et al.
Published: (2026)
by: Yao, Huanjin, et al.
Published: (2026)
EchoTracker: Advancing Myocardial Point Tracking in Echocardiography
by: Azad, Md Abulkalam, et al.
Published: (2024)
by: Azad, Md Abulkalam, et al.
Published: (2024)
LEGO: Learning and Graph-Optimized Modular Tracker for Online Multi-Object Tracking with Point Clouds
by: Zhang, Zhenrong, et al.
Published: (2023)
by: Zhang, Zhenrong, et al.
Published: (2023)
Incentivizing Tool-augmented Thinking with Images for Medical Image Analysis
by: Jiang, Yankai, et al.
Published: (2025)
by: Jiang, Yankai, et al.
Published: (2025)
Hierarchical Multi-modal Transformer for Cross-modal Long Document Classification
by: Liu, Tengfei, et al.
Published: (2024)
by: Liu, Tengfei, et al.
Published: (2024)
Prefix-Adaptive Block Diffusion for Efficient Document Recognition
by: Chai, Mingxu, et al.
Published: (2026)
by: Chai, Mingxu, et al.
Published: (2026)
Evolving High-Quality Rendering and Reconstruction in a Unified Framework with Contribution-Adaptive Regularization
by: Shen, You, et al.
Published: (2025)
by: Shen, You, et al.
Published: (2025)
LatentEdit: Adaptive Latent Control for Consistent Semantic Editing
by: Liu, Siyi, et al.
Published: (2025)
by: Liu, Siyi, et al.
Published: (2025)
Ego-Grounding for Personalized Question-Answering in Egocentric Videos
by: Xiao, Junbin, et al.
Published: (2026)
by: Xiao, Junbin, et al.
Published: (2026)
Towards Efficient Diffusion-Based Image Editing with Instant Attention Masks
by: Zou, Siyu, et al.
Published: (2024)
by: Zou, Siyu, et al.
Published: (2024)
Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts
by: Xu, Haolei, et al.
Published: (2026)
by: Xu, Haolei, et al.
Published: (2026)
Adaptive Channel Allocation for Robust Differentiable Architecture Search
by: Li, Chao, et al.
Published: (2022)
by: Li, Chao, et al.
Published: (2022)
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
by: Jung, Minjoon, et al.
Published: (2025)
by: Jung, Minjoon, et al.
Published: (2025)
Object-Level Verbalized Confidence Calibration in Vision-Language Models via Semantic Perturbation
by: Zhao, Yunpu, et al.
Published: (2025)
by: Zhao, Yunpu, et al.
Published: (2025)
Fraesormer: Learning Adaptive Sparse Transformer for Efficient Food Recognition
by: Zou, Shun, et al.
Published: (2025)
by: Zou, Shun, et al.
Published: (2025)
Mode-as-Sequence: Translating Multimodal Motion Prediction into Unified Sequential Mode Modeling
by: Zhou, Zikang, et al.
Published: (2026)
by: Zhou, Zikang, et al.
Published: (2026)
Alternative Telescopic Displacement: An Efficient Multimodal Alignment Method
by: Qin, Jiahao, et al.
Published: (2023)
by: Qin, Jiahao, et al.
Published: (2023)
Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models
by: Shi, Yuheng, et al.
Published: (2026)
by: Shi, Yuheng, et al.
Published: (2026)
Xmodel-VLM: A Simple Baseline for Multimodal Vision Language Model
by: Xu, Wanting, et al.
Published: (2024)
by: Xu, Wanting, et al.
Published: (2024)
On the Limits of Token Reduction for Efficient Unified Vision Language Training
by: Chen, Siyi, et al.
Published: (2026)
by: Chen, Siyi, et al.
Published: (2026)
PG-NeuS: Robust and Efficient Point Guidance for Multi-View Neural Surface Reconstruction
by: Zhang, Chen, et al.
Published: (2023)
by: Zhang, Chen, et al.
Published: (2023)
MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models
by: Zhao, Haozhe, et al.
Published: (2025)
by: Zhao, Haozhe, et al.
Published: (2025)
Efficient Construction of Implicit Surface Models From a Single Image for Motion Generation
by: Chu, Wei-Teng, et al.
Published: (2025)
by: Chu, Wei-Teng, et al.
Published: (2025)
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding
by: Zhang, Peng, et al.
Published: (2026)
by: Zhang, Peng, et al.
Published: (2026)
CreativeSynth: Cross-Art-Attention for Artistic Image Synthesis with Multimodal Diffusion
by: Huang, Nisha, et al.
Published: (2024)
by: Huang, Nisha, et al.
Published: (2024)
Similar Items
-
DARTer: Dynamic Adaptive Representation Tracker for Nighttime UAV Tracking
by: Li, Xuzhao, et al.
Published: (2025) -
PhyTracker: An Online Tracker for Phytoplankton
by: Yu, Yang, et al.
Published: (2024) -
Geometry OR Tracker: Universal Geometric Operating Room Tracking
by: Shao, Yihua, et al.
Published: (2026) -
A Simple Aerial Detection Baseline of Multimodal Language Models
by: Li, Qingyun, et al.
Published: (2025) -
A Simple Background Augmentation Method for Object Detection with Diffusion Model
by: Li, Yuhang, et al.
Published: (2024)