SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Alansari, Mohamad, Suryanto, Naufal, Velayudhan, Divya, Javed, Sajid, Werghi, Naoufel, Naseer, Muzammal |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking Memory Design in SAM-Based Visual Object Tracking
by: Alansari, Mohamad, et al.
Published: (2025)
by: Alansari, Mohamad, et al.
Published: (2025)
CLDTracker: A Comprehensive Language Description for Visual Tracking
by: Alansari, Mohamad, et al.
Published: (2025)
by: Alansari, Mohamad, et al.
Published: (2025)
Cytoplasmic Strings Analysis in Human Embryo Time-Lapse Videos using Deep Learning Framework
by: Sohail, Anabia, et al.
Published: (2025)
by: Sohail, Anabia, et al.
Published: (2025)
STING-BEE: Towards Vision-Language Model for Real-World X-ray Baggage Security Inspection
by: Velayudhan, Divya, et al.
Published: (2025)
by: Velayudhan, Divya, et al.
Published: (2025)
DyCON: Dynamic Uncertainty-aware Consistency and Contrastive Learning for Semi-supervised Medical Image Segmentation
by: Assefa, Maregu, et al.
Published: (2025)
by: Assefa, Maregu, et al.
Published: (2025)
Video Anomaly Detection in 10 Years: A Survey and Outlook
by: Abdalla, Moshira, et al.
Published: (2024)
by: Abdalla, Moshira, et al.
Published: (2024)
MUOT_3M: A 3 Million Frame Multimodal Underwater Benchmark and the MUTrack Tracking Method
by: Bakht, Ahsan Baidar, et al.
Published: (2026)
by: Bakht, Ahsan Baidar, et al.
Published: (2026)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
by: Wasim, Syed Talal, et al.
Published: (2023)
by: Wasim, Syed Talal, et al.
Published: (2023)
AquaticCLIP: A Vision-Language Foundation Model for Underwater Scene Analysis
by: Alawode, Basit, et al.
Published: (2025)
by: Alawode, Basit, et al.
Published: (2025)
CPLIP: Zero-Shot Learning for Histopathology with Comprehensive Vision-Language Alignment
by: Javed, Sajid, et al.
Published: (2024)
by: Javed, Sajid, et al.
Published: (2024)
Advancing Histopathology with Deep Learning Under Data Scarcity: A Decade in Review
by: Obeid, Ahmad, et al.
Published: (2024)
by: Obeid, Ahmad, et al.
Published: (2024)
Multi-Resolution Pathology-Language Pre-training Model with Text-Guided Visual Representation
by: Albastaki, Shahad, et al.
Published: (2025)
by: Albastaki, Shahad, et al.
Published: (2025)
StableMamba: Distillation-free Scaling of Large SSMs for Images and Videos
by: Suleman, Hamid, et al.
Published: (2024)
by: Suleman, Hamid, et al.
Published: (2024)
STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?
by: Li, Yun, et al.
Published: (2025)
by: Li, Yun, et al.
Published: (2025)
Video-Panda: Parameter-efficient Alignment for Encoder-free Video-Language Models
by: Yi, Jinhui, et al.
Published: (2024)
by: Yi, Jinhui, et al.
Published: (2024)
Makeup-Guided Facial Privacy Protection via Untrained Neural Network Priors
by: Shamshad, Fahad, et al.
Published: (2024)
by: Shamshad, Fahad, et al.
Published: (2024)
VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs
by: Bharadwaj, Rohit, et al.
Published: (2024)
by: Bharadwaj, Rohit, et al.
Published: (2024)
Multi-Modal Attention Networks for Enhanced Segmentation and Depth Estimation of Subsurface Defects in Pulse Thermography
by: Salah, Mohammed, et al.
Published: (2025)
by: Salah, Mohammed, et al.
Published: (2025)
Cross-Modal Self-Training: Aligning Images and Pointclouds to Learn Classification without Labels
by: Dharmasiri, Amaya, et al.
Published: (2024)
by: Dharmasiri, Amaya, et al.
Published: (2024)
VURF: A General-purpose Reasoning and Self-refinement Framework for Video Understanding
by: Mahmood, Ahmad, et al.
Published: (2024)
by: Mahmood, Ahmad, et al.
Published: (2024)
Adversarial Manhole: Challenging Monocular Depth Estimation and Semantic Segmentation Models with Patch Attack
by: Suryanto, Naufal, et al.
Published: (2024)
by: Suryanto, Naufal, et al.
Published: (2024)
RobMOT: Robust 3D Multi-Object Tracking by Observational Noise and State Estimation Drift Mitigation on LiDAR PointCloud
by: Nagy, Mohamed, et al.
Published: (2024)
by: Nagy, Mohamed, et al.
Published: (2024)
Towards Accurate State Estimation: Kalman Filter Incorporating Motion Dynamics for 3D Multi-Object Tracking
by: Nagy, Mohamed, et al.
Published: (2025)
by: Nagy, Mohamed, et al.
Published: (2025)
PromptSmooth: Certifying Robustness of Medical Vision-Language Models via Prompt Learning
by: Hussein, Noor, et al.
Published: (2024)
by: Hussein, Noor, et al.
Published: (2024)
SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs
by: Gu, Bo, et al.
Published: (2026)
by: Gu, Bo, et al.
Published: (2026)
MedContext: Learning Contextual Cues for Efficient Volumetric Medical Segmentation
by: Gani, Hanan, et al.
Published: (2024)
by: Gani, Hanan, et al.
Published: (2024)
Transformer-Based Wireless Capsule Endoscopy Bleeding Tissue Detection and Classification
by: Alawode, Basit, et al.
Published: (2024)
by: Alawode, Basit, et al.
Published: (2024)
AdaRD-key: Adaptive Relevance-Diversity Keyframe Sampling for Long-form Video understanding
by: Zhang, Xian, et al.
Published: (2025)
by: Zhang, Xian, et al.
Published: (2025)
Reinforcing Consistency in Video MLLMs with Structured Rewards
by: Quan, Yihao, et al.
Published: (2026)
by: Quan, Yihao, et al.
Published: (2026)
T2SGrid: Temporal-to-Spatial Gridification for Video Temporal Grounding
by: Guo, Chaohong, et al.
Published: (2026)
by: Guo, Chaohong, et al.
Published: (2026)
Language Guided Domain Generalized Medical Image Segmentation
by: Kunhimon, Shahina, et al.
Published: (2024)
by: Kunhimon, Shahina, et al.
Published: (2024)
SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language Models
by: Sun, Ye, et al.
Published: (2025)
by: Sun, Ye, et al.
Published: (2025)
How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
Efficient Video Face Enhancement with Enhanced Spatial-Temporal Consistency
by: Wang, Yutong, et al.
Published: (2024)
by: Wang, Yutong, et al.
Published: (2024)
Structured Video-Language Modeling with Temporal Grouping and Spatial Grounding
by: Xiong, Yuanhao, et al.
Published: (2023)
by: Xiong, Yuanhao, et al.
Published: (2023)
Composed Video Retrieval via Enriched Context and Discriminative Embeddings
by: Thawakar, Omkar, et al.
Published: (2024)
by: Thawakar, Omkar, et al.
Published: (2024)
Learning to Prompt with Text Only Supervision for Vision-Language Models
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
Multi-modal Generation via Cross-Modal In-Context Learning
by: Kumar, Amandeep, et al.
Published: (2024)
by: Kumar, Amandeep, et al.
Published: (2024)
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
by: Ouyang, Kun, et al.
Published: (2025)
by: Ouyang, Kun, et al.
Published: (2025)
Hierarchical Text-to-Vision Self Supervised Alignment for Improved Histopathology Representation Learning
by: Watawana, Hasindri, et al.
Published: (2024)
by: Watawana, Hasindri, et al.
Published: (2024)
Similar Items
-
Rethinking Memory Design in SAM-Based Visual Object Tracking
by: Alansari, Mohamad, et al.
Published: (2025) -
CLDTracker: A Comprehensive Language Description for Visual Tracking
by: Alansari, Mohamad, et al.
Published: (2025) -
Cytoplasmic Strings Analysis in Human Embryo Time-Lapse Videos using Deep Learning Framework
by: Sohail, Anabia, et al.
Published: (2025) -
STING-BEE: Towards Vision-Language Model for Real-World X-ray Baggage Security Inspection
by: Velayudhan, Divya, et al.
Published: (2025) -
DyCON: Dynamic Uncertainty-aware Consistency and Contrastive Learning for Semi-supervised Medical Image Segmentation
by: Assefa, Maregu, et al.
Published: (2025)