Rethinking Memory Design in SAM-Based Visual Object Tracking
Fuente:
arXiv
Saved in:
| Main Authors: | Alansari, Mohamad, Naseer, Muzammal, Marzouqi, Hasan Al, Werghi, Naoufel, Javed, Sajid |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMs
by: Alansari, Mohamad, et al.
Published: (2026)
by: Alansari, Mohamad, et al.
Published: (2026)
CLDTracker: A Comprehensive Language Description for Visual Tracking
by: Alansari, Mohamad, et al.
Published: (2025)
by: Alansari, Mohamad, et al.
Published: (2025)
Cytoplasmic Strings Analysis in Human Embryo Time-Lapse Videos using Deep Learning Framework
by: Sohail, Anabia, et al.
Published: (2025)
by: Sohail, Anabia, et al.
Published: (2025)
MUOT_3M: A 3 Million Frame Multimodal Underwater Benchmark and the MUTrack Tracking Method
by: Bakht, Ahsan Baidar, et al.
Published: (2026)
by: Bakht, Ahsan Baidar, et al.
Published: (2026)
Video Anomaly Detection in 10 Years: A Survey and Outlook
by: Abdalla, Moshira, et al.
Published: (2024)
by: Abdalla, Moshira, et al.
Published: (2024)
DyCON: Dynamic Uncertainty-aware Consistency and Contrastive Learning for Semi-supervised Medical Image Segmentation
by: Assefa, Maregu, et al.
Published: (2025)
by: Assefa, Maregu, et al.
Published: (2025)
AquaticCLIP: A Vision-Language Foundation Model for Underwater Scene Analysis
by: Alawode, Basit, et al.
Published: (2025)
by: Alawode, Basit, et al.
Published: (2025)
Towards Accurate State Estimation: Kalman Filter Incorporating Motion Dynamics for 3D Multi-Object Tracking
by: Nagy, Mohamed, et al.
Published: (2025)
by: Nagy, Mohamed, et al.
Published: (2025)
STING-BEE: Towards Vision-Language Model for Real-World X-ray Baggage Security Inspection
by: Velayudhan, Divya, et al.
Published: (2025)
by: Velayudhan, Divya, et al.
Published: (2025)
Multi-Resolution Pathology-Language Pre-training Model with Text-Guided Visual Representation
by: Albastaki, Shahad, et al.
Published: (2025)
by: Albastaki, Shahad, et al.
Published: (2025)
RobMOT: Robust 3D Multi-Object Tracking by Observational Noise and State Estimation Drift Mitigation on LiDAR PointCloud
by: Nagy, Mohamed, et al.
Published: (2024)
by: Nagy, Mohamed, et al.
Published: (2024)
MVTD: A Benchmark Dataset for Maritime Visual Object Tracking
by: Bakht, Ahsan Baidar, et al.
Published: (2025)
by: Bakht, Ahsan Baidar, et al.
Published: (2025)
A Distractor-Aware Memory for Visual Object Tracking with SAM2
by: Videnovic, Jovana, et al.
Published: (2024)
by: Videnovic, Jovana, et al.
Published: (2024)
Multi-Granularity Language-Guided Training for Multi-Object Tracking
by: Li, Yuhao, et al.
Published: (2024)
by: Li, Yuhao, et al.
Published: (2024)
SAMITE: Position Prompted SAM2 with Calibrated Memory for Visual Object Tracking
by: Xu, Qianxiong, et al.
Published: (2025)
by: Xu, Qianxiong, et al.
Published: (2025)
Advancing Histopathology with Deep Learning Under Data Scarcity: A Decade in Review
by: Obeid, Ahmad, et al.
Published: (2024)
by: Obeid, Ahmad, et al.
Published: (2024)
CPLIP: Zero-Shot Learning for Histopathology with Comprehensive Vision-Language Alignment
by: Javed, Sajid, et al.
Published: (2024)
by: Javed, Sajid, et al.
Published: (2024)
ObjectCompose: Evaluating Resilience of Vision-Based Models on Object-to-Background Compositional Changes
by: Malik, Hashmat Shadab, et al.
Published: (2024)
by: Malik, Hashmat Shadab, et al.
Published: (2024)
MixANT: Observation-dependent Memory Propagation for Stochastic Dense Action Anticipation
by: Wasim, Syed Talal, et al.
Published: (2025)
by: Wasim, Syed Talal, et al.
Published: (2025)
Rethinking Transformers Pre-training for Multi-Spectral Satellite Imagery
by: Noman, Mubashir, et al.
Published: (2024)
by: Noman, Mubashir, et al.
Published: (2024)
Distractor-Aware Memory-Based Visual Object Tracking
by: Videnovic, Jovana, et al.
Published: (2025)
by: Videnovic, Jovana, et al.
Published: (2025)
Efficient-SAM2: Accelerating SAM2 with Object-Aware Visual Encoding and Memory Retrieval
by: Zhang, Jing, et al.
Published: (2026)
by: Zhang, Jing, et al.
Published: (2026)
Makeup-Guided Facial Privacy Protection via Untrained Neural Network Priors
by: Shamshad, Fahad, et al.
Published: (2024)
by: Shamshad, Fahad, et al.
Published: (2024)
Enhancing Novel Object Detection via Cooperative Foundational Models
by: Bharadwaj, Rohit, et al.
Published: (2023)
by: Bharadwaj, Rohit, et al.
Published: (2023)
Y-CA-Net: A Convolutional Attention Based Network for Volumetric Medical Image Segmentation
by: Sharif, Muhammad Hamza, et al.
Published: (2024)
by: Sharif, Muhammad Hamza, et al.
Published: (2024)
Towards Evaluating the Robustness of Visual State Space Models
by: Malik, Hashmat Shadab, et al.
Published: (2024)
by: Malik, Hashmat Shadab, et al.
Published: (2024)
Language Guided Domain Generalized Medical Image Segmentation
by: Kunhimon, Shahina, et al.
Published: (2024)
by: Kunhimon, Shahina, et al.
Published: (2024)
Cross-Modal Self-Training: Aligning Images and Pointclouds to Learn Classification without Labels
by: Dharmasiri, Amaya, et al.
Published: (2024)
by: Dharmasiri, Amaya, et al.
Published: (2024)
StableMamba: Distillation-free Scaling of Large SSMs for Images and Videos
by: Suleman, Hamid, et al.
Published: (2024)
by: Suleman, Hamid, et al.
Published: (2024)
Multi-Modal Attention Networks for Enhanced Segmentation and Depth Estimation of Subsurface Defects in Pulse Thermography
by: Salah, Mohammed, et al.
Published: (2025)
by: Salah, Mohammed, et al.
Published: (2025)
ReSAM: Refine, Requery, and Reinforce: Self-Prompting Point-Supervised Segmentation for Remote Sensing Images
by: Subhani, M. Naseer
Published: (2025)
by: Subhani, M. Naseer
Published: (2025)
SAMSOD: Rethinking SAM Optimization for RGB-T Salient Object Detection
by: Liu, Zhengyi, et al.
Published: (2025)
by: Liu, Zhengyi, et al.
Published: (2025)
Probing the Efficacy of Federated Parameter-Efficient Fine-Tuning of Vision Transformers for Medical Image Classification
by: Alkhunaizi, Naif, et al.
Published: (2024)
by: Alkhunaizi, Naif, et al.
Published: (2024)
Adapting SAM 2 for Visual Object Tracking: 1st Place Solution for MMVPR Challenge Multi-Modal Tracking
by: Yang, Cheng-Yen, et al.
Published: (2025)
by: Yang, Cheng-Yen, et al.
Published: (2025)
Rethinking Backbone Design for Lightweight 3D Object Detection in LiDAR
by: Chandorkar, Adwait, et al.
Published: (2025)
by: Chandorkar, Adwait, et al.
Published: (2025)
PromptSmooth: Certifying Robustness of Medical Vision-Language Models via Prompt Learning
by: Hussein, Noor, et al.
Published: (2024)
by: Hussein, Noor, et al.
Published: (2024)
VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection
by: Liu, Chih-Chung, et al.
Published: (2026)
by: Liu, Chih-Chung, et al.
Published: (2026)
UncTrack: Reliable Visual Object Tracking with Uncertainty-Aware Prototype Memory Network
by: Yao, Siyuan, et al.
Published: (2025)
by: Yao, Siyuan, et al.
Published: (2025)
EMT: A Visual Multi-Task Benchmark Dataset for Autonomous Driving
by: Madjid, Nadya Abdel, et al.
Published: (2025)
by: Madjid, Nadya Abdel, et al.
Published: (2025)
Video-Panda: Parameter-efficient Alignment for Encoder-free Video-Language Models
by: Yi, Jinhui, et al.
Published: (2024)
by: Yi, Jinhui, et al.
Published: (2024)
Similar Items
-
SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMs
by: Alansari, Mohamad, et al.
Published: (2026) -
CLDTracker: A Comprehensive Language Description for Visual Tracking
by: Alansari, Mohamad, et al.
Published: (2025) -
Cytoplasmic Strings Analysis in Human Embryo Time-Lapse Videos using Deep Learning Framework
by: Sohail, Anabia, et al.
Published: (2025) -
MUOT_3M: A 3 Million Frame Multimodal Underwater Benchmark and the MUTrack Tracking Method
by: Bakht, Ahsan Baidar, et al.
Published: (2026) -
Video Anomaly Detection in 10 Years: A Survey and Outlook
by: Abdalla, Moshira, et al.
Published: (2024)