ReferGPT: Towards Zero-Shot Referring Multi-Object Tracking
Fuente:
arXiv
Saved in:
| Main Authors: | Chamiti, Tzoulio, Di Bella, Leandro, Munteanu, Adrian, Deligiannis, Nikos |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When Negation Is a Geometry Problem in Vision-Language Models
by: Sammani, Fawaz, et al.
Published: (2026)
by: Sammani, Fawaz, et al.
Published: (2026)
Large Models in Dialogue for Active Perception and Anomaly Detection
by: Chamiti, Tzoulio, et al.
Published: (2025)
by: Chamiti, Tzoulio, et al.
Published: (2025)
GATA2Floor: Graph attention for floor counting in street-view facades
by: Le, Ngoc Tan, et al.
Published: (2026)
by: Le, Ngoc Tan, et al.
Published: (2026)
DeepKalPose: An Enhanced Deep-Learning Kalman Filter for Temporally Consistent Monocular Vehicle Pose Estimation
by: Di Bella, Leandro, et al.
Published: (2024)
by: Di Bella, Leandro, et al.
Published: (2024)
LAM3D: Leveraging Attention for Monocular 3D Object Detection
by: Sas, Diana-Alexandra, et al.
Published: (2024)
by: Sas, Diana-Alexandra, et al.
Published: (2024)
HybridTrack: A Hybrid Approach for Robust Multi-Object Tracking
by: Di Bella, Leandro, et al.
Published: (2025)
by: Di Bella, Leandro, et al.
Published: (2025)
Cross-View Referring Multi-Object Tracking
by: Chen, Sijia, et al.
Published: (2024)
by: Chen, Sijia, et al.
Published: (2024)
Interpreting and Analysing CLIP's Zero-Shot Image Classification via Mutual Knowledge
by: Sammani, Fawaz, et al.
Published: (2024)
by: Sammani, Fawaz, et al.
Published: (2024)
MEX: Memory-efficient Approach to Referring Multi-Object Tracking
by: Tran, Huu-Thien, et al.
Published: (2025)
by: Tran, Huu-Thien, et al.
Published: (2025)
STORM: End-to-End Referring Multi-Object Tracking in Videos
by: Lu, Zijia, et al.
Published: (2026)
by: Lu, Zijia, et al.
Published: (2026)
DRMOT: A Dataset and Framework for RGBD Referring Multi-Object Tracking
by: Chen, Sijia, et al.
Published: (2026)
by: Chen, Sijia, et al.
Published: (2026)
Tell Me What to Track: Infusing Robust Language Guidance for Enhanced Referring Multi-Object Tracking
by: Huang, Wenjun, et al.
Published: (2024)
by: Huang, Wenjun, et al.
Published: (2024)
Zero-to-Hero: Zero-Shot Initialization Empowering Reference-Based Video Appearance Editing
by: Su, Tongtong, et al.
Published: (2025)
by: Su, Tongtong, et al.
Published: (2025)
FastTrackTr:Towards Fast Multi-Object Tracking with Transformers
by: Liao, Pan, et al.
Published: (2024)
by: Liao, Pan, et al.
Published: (2024)
Zero-Shot Referring Expression Comprehension via Vison-Language True/False Verification
by: Liu, Jeffrey, et al.
Published: (2025)
by: Liu, Jeffrey, et al.
Published: (2025)
Hybrid Global-Local Representation with Augmented Spatial Guidance for Zero-Shot Referring Image Segmentation
by: Liu, Ting, et al.
Published: (2025)
by: Liu, Ting, et al.
Published: (2025)
TrackTeller: Temporal Multimodal 3D Grounding for Behavior-Dependent Object References
by: Yu, Jiahong, et al.
Published: (2025)
by: Yu, Jiahong, et al.
Published: (2025)
Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes
by: Wang, Yaoting, et al.
Published: (2024)
by: Wang, Yaoting, et al.
Published: (2024)
Referring Expression Comprehension for Small Objects
by: Goto, Kanoko, et al.
Published: (2025)
by: Goto, Kanoko, et al.
Published: (2025)
ViewSAM: Learning View-aware Cross-modal Semantics for Weakly Supervised Cross-view Referring Multi-Object Tracking
by: Ge, Jiawei, et al.
Published: (2026)
by: Ge, Jiawei, et al.
Published: (2026)
VTG-GPT: Tuning-Free Zero-Shot Video Temporal Grounding with GPT
by: Xu, Yifang, et al.
Published: (2024)
by: Xu, Yifang, et al.
Published: (2024)
Training-and-Prompt-Free General Painterly Harmonization via Zero-Shot Disentenglement on Style and Content References
by: Hsiao, Teng-Fang, et al.
Published: (2024)
by: Hsiao, Teng-Fang, et al.
Published: (2024)
Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety
by: Shriram, Shashank, et al.
Published: (2025)
by: Shriram, Shashank, et al.
Published: (2025)
Efficient Multi-Object Tracking on Edge Devices via Reconstruction-Based Channel Pruning
by: Müller, Jan, et al.
Published: (2024)
by: Müller, Jan, et al.
Published: (2024)
AgentRVOS: Reasoning over Object Tracks for Zero-Shot Referring Video Object Segmentation
by: Jin, Woojeong, et al.
Published: (2026)
by: Jin, Woojeong, et al.
Published: (2026)
Mitigating Query Selection Bias in Referring Video Object Segmentation
by: Zhang, Dingwei, et al.
Published: (2025)
by: Zhang, Dingwei, et al.
Published: (2025)
InstructDET: Diversifying Referring Object Detection with Generalized Instructions
by: Dang, Ronghao, et al.
Published: (2023)
by: Dang, Ronghao, et al.
Published: (2023)
RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
by: Zhou, Enshen, et al.
Published: (2025)
by: Zhou, Enshen, et al.
Published: (2025)
DOGR: Towards Versatile Visual Document Grounding and Referring
by: Zhou, Yinan, et al.
Published: (2024)
by: Zhou, Yinan, et al.
Published: (2024)
Interpretable Open-Vocabulary Referring Object Detection with Reverse Contrast Attention
by: Juanico, Drandreb Earl O., et al.
Published: (2025)
by: Juanico, Drandreb Earl O., et al.
Published: (2025)
LLMTrack: Semantic Multi-Object Tracking with Multi-modal Large Language Models
by: Liao, Pan, et al.
Published: (2026)
by: Liao, Pan, et al.
Published: (2026)
UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
by: Liu, Ye, et al.
Published: (2025)
by: Liu, Ye, et al.
Published: (2025)
MixRI: Mixing Features of Reference Images for Novel Object Pose Estimation
by: Liu, Xinhang, et al.
Published: (2026)
by: Liu, Xinhang, et al.
Published: (2026)
DexTrack: Towards Generalizable Neural Tracking Control for Dexterous Manipulation from Human References
by: Liu, Xueyi, et al.
Published: (2025)
by: Liu, Xueyi, et al.
Published: (2025)
Vision Transformers: the threat of realistic adversarial patches
by: Cools, Kasper, et al.
Published: (2025)
by: Cools, Kasper, et al.
Published: (2025)
Awesome Multi-modal Object Tracking
by: Zhang, Chunhui, et al.
Published: (2024)
by: Zhang, Chunhui, et al.
Published: (2024)
Improving Zero-Shot Object-Level Change Detection by Incorporating Visual Correspondence
by: Nguyen, Hung Huy, et al.
Published: (2025)
by: Nguyen, Hung Huy, et al.
Published: (2025)
Self-Attention with State-Object Weighted Combination for Compositional Zero Shot Learning
by: Chang, Cheng-Hong, et al.
Published: (2025)
by: Chang, Cheng-Hong, et al.
Published: (2025)
Zero Shot Context-Based Object Segmentation using SLIP (SAM+CLIP)
by: Gundavarapu, Saaketh Koundinya, et al.
Published: (2024)
by: Gundavarapu, Saaketh Koundinya, et al.
Published: (2024)
Hybrid Discriminative Attribute-Object Embedding Network for Compositional Zero-Shot Learning
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
Similar Items
-
When Negation Is a Geometry Problem in Vision-Language Models
by: Sammani, Fawaz, et al.
Published: (2026) -
Large Models in Dialogue for Active Perception and Anomaly Detection
by: Chamiti, Tzoulio, et al.
Published: (2025) -
GATA2Floor: Graph attention for floor counting in street-view facades
by: Le, Ngoc Tan, et al.
Published: (2026) -
DeepKalPose: An Enhanced Deep-Learning Kalman Filter for Temporally Consistent Monocular Vehicle Pose Estimation
by: Di Bella, Leandro, et al.
Published: (2024) -
LAM3D: Leveraging Attention for Monocular 3D Object Detection
by: Sas, Diana-Alexandra, et al.
Published: (2024)