Towards Universal Modal Tracking with Online Dense Temporal Token Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zheng, Yaozong, Zhong, Bineng, Liang, Qihua, Zhang, Shengping, Li, Guorong, Li, Xianxian, Ji, Rongrong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ODTrack: Online Dense Temporal Token Learning for Visual Tracking
von: Zheng, Yaozong, et al.
Veröffentlicht: (2024)
von: Zheng, Yaozong, et al.
Veröffentlicht: (2024)
Less is More: Token Context-aware Learning for Object Tracking
von: Xu, Chenlong, et al.
Veröffentlicht: (2025)
von: Xu, Chenlong, et al.
Veröffentlicht: (2025)
Long-Range Feature Propagating for Natural Image Matting
von: Liu, Qinglin, et al.
Veröffentlicht: (2021)
von: Liu, Qinglin, et al.
Veröffentlicht: (2021)
Decoupled Spatio-Temporal Consistency Learning for Self-Supervised Tracking
von: Zheng, Yaozong, et al.
Veröffentlicht: (2025)
von: Zheng, Yaozong, et al.
Veröffentlicht: (2025)
UBATrack: Spatio-Temporal State Space Model for General Multi-Modal Tracking
von: Liang, Qihua, et al.
Veröffentlicht: (2026)
von: Liang, Qihua, et al.
Veröffentlicht: (2026)
Explicit Visual Prompts for Visual Object Tracking
von: Shi, Liangtao, et al.
Veröffentlicht: (2024)
von: Shi, Liangtao, et al.
Veröffentlicht: (2024)
Robust RGB-T Tracking via Learnable Visual Fourier Prompt Fine-tuning and Modality Fusion Prompt Generation
von: Yang, Hongtao, et al.
Veröffentlicht: (2025)
von: Yang, Hongtao, et al.
Veröffentlicht: (2025)
Learning to Track Instance from Single Nature Language Description
von: Zheng, Yaozong, et al.
Veröffentlicht: (2026)
von: Zheng, Yaozong, et al.
Veröffentlicht: (2026)
Boosting Self-Supervised Tracking with Contextual Prompts and Noise Learning
von: Zheng, Yaozong, et al.
Veröffentlicht: (2026)
von: Zheng, Yaozong, et al.
Veröffentlicht: (2026)
Parallel Vision Token Scheduling for Fast and Accurate Multimodal LMMs Inference
von: Zhan, Wengyi, et al.
Veröffentlicht: (2025)
von: Zhan, Wengyi, et al.
Veröffentlicht: (2025)
Similarity-Guided Layer-Adaptive Vision Transformer for UAV Tracking
von: Xue, Chaocan, et al.
Veröffentlicht: (2025)
von: Xue, Chaocan, et al.
Veröffentlicht: (2025)
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing
von: Zheng, Junjie, et al.
Veröffentlicht: (2025)
von: Zheng, Junjie, et al.
Veröffentlicht: (2025)
MambaLCT: Boosting Tracking via Long-term Context State Space Model
von: Li, Xiaohai, et al.
Veröffentlicht: (2024)
von: Li, Xiaohai, et al.
Veröffentlicht: (2024)
Dual-branch Distilled Transformer for Efficient Asymmetric UAV Tracking
von: Yang, Hongtao, et al.
Veröffentlicht: (2026)
von: Yang, Hongtao, et al.
Veröffentlicht: (2026)
Robust Tracking via Mamba-based Context-aware Token Learning
von: Xie, Jinxia, et al.
Veröffentlicht: (2024)
von: Xie, Jinxia, et al.
Veröffentlicht: (2024)
Autoregressive Queries for Adaptive Tracking with Spatio-TemporalTransformers
von: Xie, Jinxia, et al.
Veröffentlicht: (2024)
von: Xie, Jinxia, et al.
Veröffentlicht: (2024)
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos
von: Kong, Fanheng, et al.
Veröffentlicht: (2025)
von: Kong, Fanheng, et al.
Veröffentlicht: (2025)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
SpikeMba: Multi-Modal Spiking Saliency Mamba for Temporal Video Grounding
von: Li, Wenrui, et al.
Veröffentlicht: (2024)
von: Li, Wenrui, et al.
Veröffentlicht: (2024)
Cross-domain Few-shot In-context Learning for Enhancing Traffic Sign Recognition
von: Gan, Yaozong, et al.
Veröffentlicht: (2024)
von: Gan, Yaozong, et al.
Veröffentlicht: (2024)
Towards Unbiased Cross-Modal Representation Learning for Food Image-to-Recipe Retrieval
von: Wang, Qing, et al.
Veröffentlicht: (2025)
von: Wang, Qing, et al.
Veröffentlicht: (2025)
Anchoring Emotions in Text: Robust Multimodal Fusion for Mimicry Intensity Estimation
von: Zhu, Lingsi, et al.
Veröffentlicht: (2026)
von: Zhu, Lingsi, et al.
Veröffentlicht: (2026)
Hierarchical Masked Autoregressive Models with Low-Resolution Token Pivots
von: Zheng, Guangting, et al.
Veröffentlicht: (2025)
von: Zheng, Guangting, et al.
Veröffentlicht: (2025)
Multi-Modal Image Fusion via Intervention-Stable Feature Learning
von: Wang, Xue, et al.
Veröffentlicht: (2026)
von: Wang, Xue, et al.
Veröffentlicht: (2026)
Towards Flexible Evaluation for Generative Visual Question Answering
von: Ji, Huishan, et al.
Veröffentlicht: (2024)
von: Ji, Huishan, et al.
Veröffentlicht: (2024)
Exploring Phrase-Level Grounding with Text-to-Image Diffusion Model
von: Yang, Danni, et al.
Veröffentlicht: (2024)
von: Yang, Danni, et al.
Veröffentlicht: (2024)
SAM as the Guide: Mastering Pseudo-Label Refinement in Semi-Supervised Referring Expression Segmentation
von: Yang, Danni, et al.
Veröffentlicht: (2024)
von: Yang, Danni, et al.
Veröffentlicht: (2024)
2D or 3D: Who Governs Salience in VLA Models? -- Tri-Stage Token Pruning Framework with Modality Salience Awareness
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
Robust Modality-incomplete Anomaly Detection: A Modality-instructive Framework with Benchmark
von: Miao, Bingchen, et al.
Veröffentlicht: (2024)
von: Miao, Bingchen, et al.
Veröffentlicht: (2024)
Cross-domain Multi-step Thinking: Zero-shot Fine-grained Traffic Sign Recognition in the Wild
von: Gan, Yaozong, et al.
Veröffentlicht: (2024)
von: Gan, Yaozong, et al.
Veröffentlicht: (2024)
Probabilistic Temporal Masked Attention for Cross-view Online Action Detection
von: Xie, Liping, et al.
Veröffentlicht: (2025)
von: Xie, Liping, et al.
Veröffentlicht: (2025)
Learning Brain Representation with Hierarchical Visual Embeddings
von: Zheng, Jiawen, et al.
Veröffentlicht: (2026)
von: Zheng, Jiawen, et al.
Veröffentlicht: (2026)
OS-HGAdapter: Open Semantic Hypergraph Adapter for Large Language Models Assisted Entropy-Enhanced Image-Text Alignment
von: Chen, Rongjun, et al.
Veröffentlicht: (2025)
von: Chen, Rongjun, et al.
Veröffentlicht: (2025)
Towards Robust and Realible Multimodal Misinformation Recognition with Incomplete Modality
von: Zhou, Hengyang, et al.
Veröffentlicht: (2025)
von: Zhou, Hengyang, et al.
Veröffentlicht: (2025)
Context-aware TFL: A Universal Context-aware Contrastive Learning Framework for Temporal Forgery Localization
von: Yin, Qilin, et al.
Veröffentlicht: (2025)
von: Yin, Qilin, et al.
Veröffentlicht: (2025)
M3FAS: An Accurate and Robust MultiModal Mobile Face Anti-Spoofing System
von: Kong, Chenqi, et al.
Veröffentlicht: (2023)
von: Kong, Chenqi, et al.
Veröffentlicht: (2023)
Tracking and Segmenting Anything in Any Modality
von: Zhang, Tianlu, et al.
Veröffentlicht: (2025)
von: Zhang, Tianlu, et al.
Veröffentlicht: (2025)
CoNav: Collaborative Cross-Modal Reasoning for Embodied Navigation
von: Hao, Haihong, et al.
Veröffentlicht: (2025)
von: Hao, Haihong, et al.
Veröffentlicht: (2025)
Pursuing Temporal-Consistent Video Virtual Try-On via Dynamic Pose Interaction
von: Li, Dong, et al.
Veröffentlicht: (2025)
von: Li, Dong, et al.
Veröffentlicht: (2025)
ROI-Guided Point Cloud Geometry Compression Towards Human and Machine Vision
von: Liang, Xie, et al.
Veröffentlicht: (2025)
von: Liang, Xie, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ODTrack: Online Dense Temporal Token Learning for Visual Tracking
von: Zheng, Yaozong, et al.
Veröffentlicht: (2024) -
Less is More: Token Context-aware Learning for Object Tracking
von: Xu, Chenlong, et al.
Veröffentlicht: (2025) -
Long-Range Feature Propagating for Natural Image Matting
von: Liu, Qinglin, et al.
Veröffentlicht: (2021) -
Decoupled Spatio-Temporal Consistency Learning for Self-Supervised Tracking
von: Zheng, Yaozong, et al.
Veröffentlicht: (2025) -
UBATrack: Spatio-Temporal State Space Model for General Multi-Modal Tracking
von: Liang, Qihua, et al.
Veröffentlicht: (2026)