STORM: Segment, Track, and Object Re-Localization from a Single Image
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Deng, Yu, Cao, Teng, Shindo, Hikaru, Delfosse, Quentin, Xue, Jiahong, Kersting, Kristian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Boosting Object Representation Learning via Motion and Object Continuity
von: Delfosse, Quentin, et al.
Veröffentlicht: (2022)
von: Delfosse, Quentin, et al.
Veröffentlicht: (2022)
OCAtari: Object-Centric Atari 2600 Reinforcement Learning Environments
von: Delfosse, Quentin, et al.
Veröffentlicht: (2023)
von: Delfosse, Quentin, et al.
Veröffentlicht: (2023)
Learning Differentiable Logic Programs for Abstract Visual Reasoning
von: Shindo, Hikaru, et al.
Veröffentlicht: (2023)
von: Shindo, Hikaru, et al.
Veröffentlicht: (2023)
DeiSAM: Segment Anything with Deictic Prompting
von: Shindo, Hikaru, et al.
Veröffentlicht: (2024)
von: Shindo, Hikaru, et al.
Veröffentlicht: (2024)
V-LoL: A Diagnostic Dataset for Visual Logical Learning
von: Helff, Lukas, et al.
Veröffentlicht: (2023)
von: Helff, Lukas, et al.
Veröffentlicht: (2023)
SuperPose: Improved 6D Pose Estimation with Robust Tracking and Mask-Free Initialization
von: Deng, Yu, et al.
Veröffentlicht: (2024)
von: Deng, Yu, et al.
Veröffentlicht: (2024)
ART: Adaptive Relation Tuning for Generalized Relation Prediction
von: Sudhakaran, Gopika, et al.
Veröffentlicht: (2025)
von: Sudhakaran, Gopika, et al.
Veröffentlicht: (2025)
Pix2Code: Learning to Compose Neural Visual Concepts as Programs
von: Wüst, Antonia, et al.
Veröffentlicht: (2024)
von: Wüst, Antonia, et al.
Veröffentlicht: (2024)
STORM: End-to-End Referring Multi-Object Tracking in Videos
von: Lu, Zijia, et al.
Veröffentlicht: (2026)
von: Lu, Zijia, et al.
Veröffentlicht: (2026)
GRAIL: Autonomous Concept Grounding for Neuro-Symbolic Reinforcement Learning
von: Shindo, Hikaru, et al.
Veröffentlicht: (2026)
von: Shindo, Hikaru, et al.
Veröffentlicht: (2026)
Image Coding for Machines with Edge Information Learning Using Segment Anything
von: Shindo, Takahiro, et al.
Veröffentlicht: (2024)
von: Shindo, Takahiro, et al.
Veröffentlicht: (2024)
TrackTeller: Temporal Multimodal 3D Grounding for Behavior-Dependent Object References
von: Yu, Jiahong, et al.
Veröffentlicht: (2025)
von: Yu, Jiahong, et al.
Veröffentlicht: (2025)
Joint Counting, Detection and Re-Identification for Multi-Object Tracking
von: Ren, Weihong, et al.
Veröffentlicht: (2022)
von: Ren, Weihong, et al.
Veröffentlicht: (2022)
Enhancing Video Object Segmentation in TrackRAD Using XMem Memory Network
von: Deng, Pengchao, et al.
Veröffentlicht: (2025)
von: Deng, Pengchao, et al.
Veröffentlicht: (2025)
A Simple Video Segmenter by Tracking Objects Along Axial Trajectories
von: He, Ju, et al.
Veröffentlicht: (2023)
von: He, Ju, et al.
Veröffentlicht: (2023)
ClickTrack: Towards Real-time Interactive Single Object Tracking
von: Wang, Kuiran, et al.
Veröffentlicht: (2024)
von: Wang, Kuiran, et al.
Veröffentlicht: (2024)
BlendRL: A Framework for Merging Symbolic and Neural Policy Learning
von: Shindo, Hikaru, et al.
Veröffentlicht: (2024)
von: Shindo, Hikaru, et al.
Veröffentlicht: (2024)
Temporal-Enhanced Multimodal Transformer for Referring Multi-Object Tracking and Segmentation
von: Xiao, Changcheng, et al.
Veröffentlicht: (2024)
von: Xiao, Changcheng, et al.
Veröffentlicht: (2024)
Instance-Level Moving Object Segmentation from a Single Image with Events
von: Wan, Zhexiong, et al.
Veröffentlicht: (2025)
von: Wan, Zhexiong, et al.
Veröffentlicht: (2025)
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models
von: Liang, Yiming, et al.
Veröffentlicht: (2026)
von: Liang, Yiming, et al.
Veröffentlicht: (2026)
Segment Anything for Video: A Comprehensive Review of Video Object Segmentation and Tracking from Past to Future
von: Xu, Guoping, et al.
Veröffentlicht: (2025)
von: Xu, Guoping, et al.
Veröffentlicht: (2025)
An Effective Solution for the CVPR 2026 8th UG2+ Challenge Track 3: Dynamic Object Segmentation in Turbulence
von: Li, Hongzhen, et al.
Veröffentlicht: (2026)
von: Li, Hongzhen, et al.
Veröffentlicht: (2026)
P2Object: Single Point Supervised Object Detection and Instance Segmentation
von: Chen, Pengfei, et al.
Veröffentlicht: (2025)
von: Chen, Pengfei, et al.
Veröffentlicht: (2025)
STORM: Token-Efficient Long Video Understanding for Multimodal LLMs
von: Jiang, Jindong, et al.
Veröffentlicht: (2025)
von: Jiang, Jindong, et al.
Veröffentlicht: (2025)
No Safe Dose: How Training Data Drives Unsafe Image Generation
von: Friedrich, Felix, et al.
Veröffentlicht: (2026)
von: Friedrich, Felix, et al.
Veröffentlicht: (2026)
STORM: Strategic Orchestration of Modalities for Rare Event Classification
von: Kamboj, Payal, et al.
Veröffentlicht: (2024)
von: Kamboj, Payal, et al.
Veröffentlicht: (2024)
EAFormer: Scene Text Segmentation with Edge-Aware Transformers
von: Yu, Haiyang, et al.
Veröffentlicht: (2024)
von: Yu, Haiyang, et al.
Veröffentlicht: (2024)
History-Aware Transformation of ReID Features for Multiple Object Tracking
von: Gao, Ruopeng, et al.
Veröffentlicht: (2025)
von: Gao, Ruopeng, et al.
Veröffentlicht: (2025)
Beyond Traditional Single Object Tracking: A Survey
von: Abdelaziz, Omar, et al.
Veröffentlicht: (2024)
von: Abdelaziz, Omar, et al.
Veröffentlicht: (2024)
Single-Model and Any-Modality for Video Object Tracking
von: Wu, Zongwei, et al.
Veröffentlicht: (2023)
von: Wu, Zongwei, et al.
Veröffentlicht: (2023)
Stylistic-STORM (ST-STORM) : Perceiving the Semantic Nature of Appearance
von: Ouattara, Hamed, et al.
Veröffentlicht: (2026)
von: Ouattara, Hamed, et al.
Veröffentlicht: (2026)
Beyond Euclidean Prototypes: Spectral Disentanglement and Geodesic Matching for Few-Shot Medical Image Segmentation
von: Jia, Penghao, et al.
Veröffentlicht: (2026)
von: Jia, Penghao, et al.
Veröffentlicht: (2026)
360VOTS: Visual Object Tracking and Segmentation in Omnidirectional Videos
von: Xu, Yinzhe, et al.
Veröffentlicht: (2024)
von: Xu, Yinzhe, et al.
Veröffentlicht: (2024)
SAM3-Adapter: Efficient Adaptation of Segment Anything 3 for Camouflage Object Segmentation, Shadow Detection, and Medical Image Segmentation
von: Chen, Tianrun, et al.
Veröffentlicht: (2025)
von: Chen, Tianrun, et al.
Veröffentlicht: (2025)
FlowTrack: Point-level Flow Network for 3D Single Object Tracking
von: Li, Shuo, et al.
Veröffentlicht: (2024)
von: Li, Shuo, et al.
Veröffentlicht: (2024)
VPIT: Real-time Embedded Single Object 3D Tracking Using Voxel Pseudo Images
von: Oleksiienko, Illia, et al.
Veröffentlicht: (2022)
von: Oleksiienko, Illia, et al.
Veröffentlicht: (2022)
METok: Multi-Stage Event-based Token Compression for Efficient Long Video Understanding
von: Wang, Mengyue, et al.
Veröffentlicht: (2025)
von: Wang, Mengyue, et al.
Veröffentlicht: (2025)
Multi-Object Tracking by Hierarchical Visual Representations
von: Cao, Jinkun, et al.
Veröffentlicht: (2024)
von: Cao, Jinkun, et al.
Veröffentlicht: (2024)
Segment Anything with Motion, Geometry, and Semantic Adaptation for Complex Nonlinear Visual Object Tracking
von: Zhu, Deyi, et al.
Veröffentlicht: (2026)
von: Zhu, Deyi, et al.
Veröffentlicht: (2026)
MSITrack: A Challenging Benchmark for Multispectral Single Object Tracking
von: Feng, Tao, et al.
Veröffentlicht: (2025)
von: Feng, Tao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Boosting Object Representation Learning via Motion and Object Continuity
von: Delfosse, Quentin, et al.
Veröffentlicht: (2022) -
OCAtari: Object-Centric Atari 2600 Reinforcement Learning Environments
von: Delfosse, Quentin, et al.
Veröffentlicht: (2023) -
Learning Differentiable Logic Programs for Abstract Visual Reasoning
von: Shindo, Hikaru, et al.
Veröffentlicht: (2023) -
DeiSAM: Segment Anything with Deictic Prompting
von: Shindo, Hikaru, et al.
Veröffentlicht: (2024) -
V-LoL: A Diagnostic Dataset for Visual Logical Learning
von: Helff, Lukas, et al.
Veröffentlicht: (2023)