STORM: End-to-End Referring Multi-Object Tracking in Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Lu, Zijia, Yi, Jingru, Wang, Jue, Chen, Yuxiao, Chen, Junwen, Li, Xinyu, Modolo, Davide |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Supervised Multi-Object Tracking with Path Consistency
by: Lu, Zijia, et al.
Published: (2024)
by: Lu, Zijia, et al.
Published: (2024)
Cross-View Referring Multi-Object Tracking
by: Chen, Sijia, et al.
Published: (2024)
by: Chen, Sijia, et al.
Published: (2024)
Learning Compact Video Representations for Efficient Long-form Video Understanding in Large Multimodal Models
by: Chen, Yuxiao, et al.
Published: (2026)
by: Chen, Yuxiao, et al.
Published: (2026)
Adversarial Flow Matching for Imperceptible Attacks on End-to-End Autonomous Driving
by: Zeng, Xinyu, et al.
Published: (2026)
by: Zeng, Xinyu, et al.
Published: (2026)
Towards End-to-End Neuromorphic Event-based 3D Object Reconstruction Without Physical Priors
by: Xu, Chuanzhi, et al.
Published: (2025)
by: Xu, Chuanzhi, et al.
Published: (2025)
End-To-End Underwater Video Enhancement: Dataset and Model
by: Du, Dazhao, et al.
Published: (2024)
by: Du, Dazhao, et al.
Published: (2024)
Beyond Hungarian: Match-Free Supervision for End-to-End Object Detection
by: Qiu, Shoumeng, et al.
Published: (2026)
by: Qiu, Shoumeng, et al.
Published: (2026)
ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation
by: Zhang, Mengchen, et al.
Published: (2025)
by: Zhang, Mengchen, et al.
Published: (2025)
EVA: Efficient Reinforcement Learning for End-to-End Video Agent
by: Zhang, Yaolun, et al.
Published: (2026)
by: Zhang, Yaolun, et al.
Published: (2026)
End-to-End Multi-Modal Diffusion Mamba
by: Lu, Chunhao, et al.
Published: (2025)
by: Lu, Chunhao, et al.
Published: (2025)
X-World: Controllable Ego-Centric Multi-Camera World Models for Scalable End-to-End Driving
by: Zheng, Chaoda, et al.
Published: (2026)
by: Zheng, Chaoda, et al.
Published: (2026)
DRMOT: A Dataset and Framework for RGBD Referring Multi-Object Tracking
by: Chen, Sijia, et al.
Published: (2026)
by: Chen, Sijia, et al.
Published: (2026)
S.T.A.R.-Track: Latent Motion Models for End-to-End 3D Object Tracking with Adaptive Spatio-Temporal Appearance Representations
by: Doll, Simon, et al.
Published: (2023)
by: Doll, Simon, et al.
Published: (2023)
EVC-MF: End-to-end Video Captioning Network with Multi-scale Features
by: Niu, Tian-Zi, et al.
Published: (2024)
by: Niu, Tian-Zi, et al.
Published: (2024)
ReferGPT: Towards Zero-Shot Referring Multi-Object Tracking
by: Chamiti, Tzoulio, et al.
Published: (2025)
by: Chamiti, Tzoulio, et al.
Published: (2025)
DocPTBench: Benchmarking End-to-End Photographed Document Parsing and Translation
by: Du, Yongkun, et al.
Published: (2025)
by: Du, Yongkun, et al.
Published: (2025)
Tell Me What to Track: Infusing Robust Language Guidance for Enhanced Referring Multi-Object Tracking
by: Huang, Wenjun, et al.
Published: (2024)
by: Huang, Wenjun, et al.
Published: (2024)
Multi-granularity Contrastive Cross-modal Collaborative Generation for End-to-End Long-term Video Question Answering
by: Yu, Ting, et al.
Published: (2024)
by: Yu, Ting, et al.
Published: (2024)
ContourFormer: Real-Time Contour-Based End-to-End Instance Segmentation Transformer
by: Yao, Weiwei, et al.
Published: (2025)
by: Yao, Weiwei, et al.
Published: (2025)
Guiding Attention in End-to-End Driving Models
by: Porres, Diego, et al.
Published: (2024)
by: Porres, Diego, et al.
Published: (2024)
DART: An Automated End-to-End Object Detection Pipeline with Data Diversification, Open-Vocabulary Bounding Box Annotation, Pseudo-Label Review, and Model Training
by: Xin, Chen, et al.
Published: (2024)
by: Xin, Chen, et al.
Published: (2024)
YOLOv10: Real-Time End-to-End Object Detection
by: Wang, Ao, et al.
Published: (2024)
by: Wang, Ao, et al.
Published: (2024)
MEX: Memory-efficient Approach to Referring Multi-Object Tracking
by: Tran, Huu-Thien, et al.
Published: (2025)
by: Tran, Huu-Thien, et al.
Published: (2025)
YOLO26: An Analysis of NMS-Free End to End Framework for Real-Time Object Detection
by: Chakrabarty, Sudip
Published: (2026)
by: Chakrabarty, Sudip
Published: (2026)
A Paradigm Shift: Fully End-to-End Training for Temporal Sentence Grounding in Videos
by: He, Allen, et al.
Published: (2026)
by: He, Allen, et al.
Published: (2026)
MITracker: Multi-View Integration for Visual Object Tracking
by: Xu, Mengjie, et al.
Published: (2025)
by: Xu, Mengjie, et al.
Published: (2025)
Tracking by Detection and Query: An Efficient End-to-End Framework for Multi-Object Tracking
by: Jia, Shukun, et al.
Published: (2024)
by: Jia, Shukun, et al.
Published: (2024)
Hint-AD: Holistically Aligned Interpretability in End-to-End Autonomous Driving
by: Ding, Kairui, et al.
Published: (2024)
by: Ding, Kairui, et al.
Published: (2024)
ICDAR 2025 Competition on End-to-End Document Image Machine Translation Towards Complex Layouts
by: Zhang, Yaping, et al.
Published: (2026)
by: Zhang, Yaping, et al.
Published: (2026)
iPad: Iterative Proposal-centric End-to-End Autonomous Driving
by: Guo, Ke, et al.
Published: (2025)
by: Guo, Ke, et al.
Published: (2025)
Training Multi-Image Vision Agents via End2End Reinforcement Learning
by: Dong, Chengqi, et al.
Published: (2025)
by: Dong, Chengqi, et al.
Published: (2025)
Feature Corrective Transfer Learning: End-to-End Solutions to Object Detection in Non-Ideal Visual Conditions
by: Wei, Chuheng, et al.
Published: (2024)
by: Wei, Chuheng, et al.
Published: (2024)
EA-RAS: Towards Efficient and Accurate End-to-End Reconstruction of Anatomical Skeleton
by: Peng, Zhiheng, et al.
Published: (2024)
by: Peng, Zhiheng, et al.
Published: (2024)
FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token Pruning
by: Cao, Jiajun, et al.
Published: (2025)
by: Cao, Jiajun, et al.
Published: (2025)
Without Paired Labeled Data: End-to-End Self-Supervised Learning for Drone-view Geo-Localization
by: Chen, Zhongwei, et al.
Published: (2025)
by: Chen, Zhongwei, et al.
Published: (2025)
M$^3$-VOS: Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentation
by: Chen, Zixuan, et al.
Published: (2024)
by: Chen, Zixuan, et al.
Published: (2024)
ParkingScenes: A Structured Dataset for End-to-End Autonomous Parking in Simulation Scenes
by: Chen, Haonan, et al.
Published: (2026)
by: Chen, Haonan, et al.
Published: (2026)
CLOVER: Closed-Loop Value Estimation and Ranking for End-to-End Autonomous Driving Planning
by: Ang, Sining, et al.
Published: (2026)
by: Ang, Sining, et al.
Published: (2026)
SMTrack: End-to-End Trained Spiking Neural Networks for Multi-Object Tracking in RGB Videos
by: Zhong, Pengzhi, et al.
Published: (2025)
by: Zhong, Pengzhi, et al.
Published: (2025)
RainFusion: Adaptive Video Generation Acceleration via Multi-Dimensional Visual Redundancy
by: Chen, Aiyue, et al.
Published: (2025)
by: Chen, Aiyue, et al.
Published: (2025)
Similar Items
-
Self-Supervised Multi-Object Tracking with Path Consistency
by: Lu, Zijia, et al.
Published: (2024) -
Cross-View Referring Multi-Object Tracking
by: Chen, Sijia, et al.
Published: (2024) -
Learning Compact Video Representations for Efficient Long-form Video Understanding in Large Multimodal Models
by: Chen, Yuxiao, et al.
Published: (2026) -
Adversarial Flow Matching for Imperceptible Attacks on End-to-End Autonomous Driving
by: Zeng, Xinyu, et al.
Published: (2026) -
Towards End-to-End Neuromorphic Event-based 3D Object Reconstruction Without Physical Priors
by: Xu, Chuanzhi, et al.
Published: (2025)