Unifying Visual and Vision-Language Tracking via Contrastive Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Yinchao, Tang, Yuyang, Yang, Wenfei, Zhang, Tianzhu, Zhang, Jinpeng, Kang, Mengxue |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UniSOT: A Unified Framework for Multi-Modality Single Object Tracking
by: Ma, Yinchao, et al.
Published: (2025)
by: Ma, Yinchao, et al.
Published: (2025)
SMTrack: State-Aware Mamba for Efficient Temporal Modeling in Visual Tracking
by: Ma, Yinchao, et al.
Published: (2026)
by: Ma, Yinchao, et al.
Published: (2026)
Multi-modal Attribute Prompting for Vision-Language Models
by: Liu, Xin, et al.
Published: (2024)
by: Liu, Xin, et al.
Published: (2024)
Learning Shape-Independent Transformation via Spherical Representations for Category-Level Object Pose Estimation
by: Ren, Huan, et al.
Published: (2025)
by: Ren, Huan, et al.
Published: (2025)
Instance-Adaptive and Geometric-Aware Keypoint Learning for Category-Level 6D Object Pose Estimation
by: Lin, Xiao, et al.
Published: (2024)
by: Lin, Xiao, et al.
Published: (2024)
StruMamba3D: Exploring Structural Mamba for Self-supervised Point Cloud Representation Learning
by: Wang, Chuxin, et al.
Published: (2025)
by: Wang, Chuxin, et al.
Published: (2025)
Exploring Semantic Masked Autoencoder for Self-supervised Point Cloud Understanding
by: Zha, Yixin, et al.
Published: (2025)
by: Zha, Yixin, et al.
Published: (2025)
Structure-Aware Correspondence Learning for Relative Pose Estimation
by: Chen, Yihan, et al.
Published: (2025)
by: Chen, Yihan, et al.
Published: (2025)
ComPose: A Unified Completion-Pose Framework for Robust Category-Level Object Pose Estimation
by: Ren, Huan, et al.
Published: (2026)
by: Ren, Huan, et al.
Published: (2026)
State Space Model Meets Transformer: A New Paradigm for 3D Object Detection
by: Wang, Chuxin, et al.
Published: (2025)
by: Wang, Chuxin, et al.
Published: (2025)
Exploring a Unified Vision-Centric Contrastive Alternatives on Multi-Modal Web Documents
by: Lin, Yiqi, et al.
Published: (2025)
by: Lin, Yiqi, et al.
Published: (2025)
Linear Differential Vision Transformer: Learning Visual Contrasts via Pairwise Differentials
by: Pu, Yifan, et al.
Published: (2025)
by: Pu, Yifan, et al.
Published: (2025)
Plane2Depth: Hierarchical Adaptive Plane Guidance for Monocular Depth Estimation
by: Liu, Li, et al.
Published: (2024)
by: Liu, Li, et al.
Published: (2024)
Unified Multimodal Coherent Field: Synchronous Semantic-Spatial-Vision Fusion for Brain Tumor Segmentation
by: Zhang, Mingda, et al.
Published: (2025)
by: Zhang, Mingda, et al.
Published: (2025)
C3L: Content Correlated Vision-Language Instruction Tuning Data Generation via Contrastive Learning
by: Ma, Ji, et al.
Published: (2024)
by: Ma, Ji, et al.
Published: (2024)
VPTracker: Global Vision-Language Tracking via Visual Prompt
by: Wang, Jingchao, et al.
Published: (2025)
by: Wang, Jingchao, et al.
Published: (2025)
Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision
by: Wei, Zhixiang, et al.
Published: (2026)
by: Wei, Zhixiang, et al.
Published: (2026)
DN-4DGS: Denoised Deformable Network with Temporal-Spatial Aggregation for Dynamic Scene Rendering
by: Lu, Jiahao, et al.
Published: (2024)
by: Lu, Jiahao, et al.
Published: (2024)
CFTrack: Enhancing Lightweight Visual Tracking through Contrastive Learning and Feature Matching
by: Liang, Juntao, et al.
Published: (2025)
by: Liang, Juntao, et al.
Published: (2025)
Unifying Graph Contrastive Learning via Graph Message Augmentation
by: Zhang, Ziyan, et al.
Published: (2024)
by: Zhang, Ziyan, et al.
Published: (2024)
Ultrasound Vision-Language Alignment via Contrastive Learning
by: Lyu, Zhuoyang, et al.
Published: (2026)
by: Lyu, Zhuoyang, et al.
Published: (2026)
Unified Sequence-to-Sequence Learning for Single- and Multi-Modal Visual Object Tracking
by: Chen, Xin, et al.
Published: (2023)
by: Chen, Xin, et al.
Published: (2023)
BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration
by: Li, Zhaoyang, et al.
Published: (2025)
by: Li, Zhaoyang, et al.
Published: (2025)
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning
by: Yang, Jie, et al.
Published: (2025)
by: Yang, Jie, et al.
Published: (2025)
UAUTrack: Towards Unified Multimodal Anti-UAV Visual Tracking
by: Ren, Qionglin, et al.
Published: (2025)
by: Ren, Qionglin, et al.
Published: (2025)
ACTRESS: Active Retraining for Semi-supervised Visual Grounding
by: Kang, Weitai, et al.
Published: (2024)
by: Kang, Weitai, et al.
Published: (2024)
Frequency Domain Modality-invariant Feature Learning for Visible-infrared Person Re-Identification
by: Li, Yulin, et al.
Published: (2024)
by: Li, Yulin, et al.
Published: (2024)
Pamba: Enhancing Global Interaction in Point Clouds via State Space Model
by: Li, Zhuoyuan, et al.
Published: (2024)
by: Li, Zhuoyuan, et al.
Published: (2024)
Bridging Visual Representation and Reinforcement Learning from Verifiable Rewards in Large Vision-Language Models
by: Han, Yuhang, et al.
Published: (2026)
by: Han, Yuhang, et al.
Published: (2026)
PVLM: Parsing-Aware Vision Language Model with Dynamic Contrastive Learning for Zero-Shot Deepfake Attribution
by: Zhang, Yaning, et al.
Published: (2025)
by: Zhang, Yaning, et al.
Published: (2025)
SignVTCL: Multi-Modal Continuous Sign Language Recognition Enhanced by Visual-Textual Contrastive Learning
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
VL-UniTrack: A Unified Framework with Visual-Language Prompts for UAV-Ground Visual Tracking
by: Xu, Boyue, et al.
Published: (2026)
by: Xu, Boyue, et al.
Published: (2026)
Unified Language-Vision Pretraining in LLM with Dynamic Discrete Visual Tokenization
by: Jin, Yang, et al.
Published: (2023)
by: Jin, Yang, et al.
Published: (2023)
EF-3DGS: Event-Aided Free-Trajectory 3D Gaussian Splatting
by: Liao, Bohao, et al.
Published: (2024)
by: Liao, Bohao, et al.
Published: (2024)
Unifying 3D Vision-Language Understanding via Promptable Queries
by: Zhu, Ziyu, et al.
Published: (2024)
by: Zhu, Ziyu, et al.
Published: (2024)
BoostAdapter: Improving Vision-Language Test-Time Adaptation via Regional Bootstrapping
by: Zhang, Taolin, et al.
Published: (2024)
by: Zhang, Taolin, et al.
Published: (2024)
All in One: Exploring Unified Vision-Language Tracking with Multi-Modal Alignment
by: Zhang, Chunhui, et al.
Published: (2023)
by: Zhang, Chunhui, et al.
Published: (2023)
UTPTrack: Towards Simple and Unified Token Pruning for Visual Tracking
by: Wu, Hao, et al.
Published: (2026)
by: Wu, Hao, et al.
Published: (2026)
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues
by: Feng, X., et al.
Published: (2024)
by: Feng, X., et al.
Published: (2024)
Unified Multimodal Visual Tracking with Dual Mixture-of-Experts
by: Hong, Lingyi, et al.
Published: (2026)
by: Hong, Lingyi, et al.
Published: (2026)
Similar Items
-
UniSOT: A Unified Framework for Multi-Modality Single Object Tracking
by: Ma, Yinchao, et al.
Published: (2025) -
SMTrack: State-Aware Mamba for Efficient Temporal Modeling in Visual Tracking
by: Ma, Yinchao, et al.
Published: (2026) -
Multi-modal Attribute Prompting for Vision-Language Models
by: Liu, Xin, et al.
Published: (2024) -
Learning Shape-Independent Transformation via Spherical Representations for Category-Level Object Pose Estimation
by: Ren, Huan, et al.
Published: (2025) -
Instance-Adaptive and Geometric-Aware Keypoint Learning for Category-Level 6D Object Pose Estimation
by: Lin, Xiao, et al.
Published: (2024)