Open-World Object Counting in Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Amini-Naieni, Niki, Zisserman, Andrew |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CountGD++: Generalized Prompting for Open-World Counting
by: Amini-Naieni, Niki, et al.
Published: (2025)
by: Amini-Naieni, Niki, et al.
Published: (2025)
CountGD: Multi-Modal Open-World Counting
by: Amini-Naieni, Niki, et al.
Published: (2024)
by: Amini-Naieni, Niki, et al.
Published: (2024)
The MixCount Dataset: Bridging the Data Gap for Open-Vocabulary Object Counting
by: Dumery, Corentin, et al.
Published: (2026)
by: Dumery, Corentin, et al.
Published: (2026)
Unveiling Ontological Commitment in Multi-Modal Foundation Models
by: Keser, Mert, et al.
Published: (2024)
by: Keser, Mert, et al.
Published: (2024)
OVR: A Dataset for Open Vocabulary Temporal Repetition Counting in Videos
by: Dwibedi, Debidatta, et al.
Published: (2024)
by: Dwibedi, Debidatta, et al.
Published: (2024)
A Short Note on Evaluating RepNet for Temporal Repetition Counting in Videos
by: Dwibedi, Debidatta, et al.
Published: (2024)
by: Dwibedi, Debidatta, et al.
Published: (2024)
Instant Uncertainty Calibration of NeRFs Using a Meta-Calibrator
by: Amini-Naieni, Niki, et al.
Published: (2023)
by: Amini-Naieni, Niki, et al.
Published: (2023)
New keypoint-based approach for recognising British Sign Language (BSL) from sequences
by: Deb, Oishi, et al.
Published: (2024)
by: Deb, Oishi, et al.
Published: (2024)
VOVTrack: Exploring the Potentiality in Videos for Open-Vocabulary Object Tracking
by: Qian, Zekun, et al.
Published: (2024)
by: Qian, Zekun, et al.
Published: (2024)
From Open Vocabulary to Open World: Teaching Vision Language Models to Detect Novel Objects
by: Li, Zizhao, et al.
Published: (2024)
by: Li, Zizhao, et al.
Published: (2024)
Dream2Flow: Bridging Video Generation and Open-World Manipulation with 3D Object Flow
by: Dharmarajan, Karthik, et al.
Published: (2025)
by: Dharmarajan, Karthik, et al.
Published: (2025)
Semantic Compression of 3D Objects for Open and Collaborative Virtual Worlds
by: Dotzel, Jordan, et al.
Published: (2025)
by: Dotzel, Jordan, et al.
Published: (2025)
DualMem: Bypassing the Objectness Bottleneck for Calibrated Unknown-Stream Filtering in Open-World Object Detection
by: Xiao, Yingjun, et al.
Published: (2026)
by: Xiao, Yingjun, et al.
Published: (2026)
3D-Aware Instance Segmentation and Tracking in Egocentric Videos
by: Bhalgat, Yash, et al.
Published: (2024)
by: Bhalgat, Yash, et al.
Published: (2024)
GroundCount: Grounding Vision-Language Models with Object Detection for Mitigating Counting Hallucinations
by: Chen, Boyuan, et al.
Published: (2026)
by: Chen, Boyuan, et al.
Published: (2026)
OV-DQUO: Open-Vocabulary DETR with Denoising Text Query Training and Open-World Unknown Objects Supervision
by: Wang, Junjie, et al.
Published: (2024)
by: Wang, Junjie, et al.
Published: (2024)
Learning from One Continuous Video Stream
by: Carreira, João, et al.
Published: (2023)
by: Carreira, João, et al.
Published: (2023)
OW-VISCapTor: Abstractors for Open-World Video Instance Segmentation and Captioning
by: Choudhuri, Anwesa, et al.
Published: (2024)
by: Choudhuri, Anwesa, et al.
Published: (2024)
CountFormer: A Transformer Framework for Learning Visual Repetition and Structure in Class-Agnostic Object Counting
by: Hossain, Md Tanvir, et al.
Published: (2025)
by: Hossain, Md Tanvir, et al.
Published: (2025)
MambaLoc: Efficient Camera Localisation via State Space Model
by: Wang, Jialu, et al.
Published: (2024)
by: Wang, Jialu, et al.
Published: (2024)
Benchmarking Vision Foundation Models for Input Monitoring in Autonomous Driving
by: Keser, Mert, et al.
Published: (2025)
by: Keser, Mert, et al.
Published: (2025)
Self-supervised Learning on Camera Trap Footage Yields a Strong Universal Face Embedder
by: Iashin, Vladimir, et al.
Published: (2025)
by: Iashin, Vladimir, et al.
Published: (2025)
WoMAP: World Models For Embodied Open-Vocabulary Object Localization
by: Yin, Tenny, et al.
Published: (2025)
by: Yin, Tenny, et al.
Published: (2025)
Mutually-Aware Feature Learning for Few-Shot Object Counting
by: Jeon, Yerim, et al.
Published: (2024)
by: Jeon, Yerim, et al.
Published: (2024)
Open-CRB: Towards Open World Active Learning for 3D Object Detection
by: Chen, Zhuoxiao, et al.
Published: (2023)
by: Chen, Zhuoxiao, et al.
Published: (2023)
Seed Kernel Counting using Domain Randomization and Object Tracking Neural Networks
by: Margapuri, Venkat, et al.
Published: (2023)
by: Margapuri, Venkat, et al.
Published: (2023)
World Models for Learning Dexterous Hand-Object Interactions from Human Videos
by: Goswami, Raktim Gautam, et al.
Published: (2025)
by: Goswami, Raktim Gautam, et al.
Published: (2025)
PhysWorld: From Real Videos to World Models of Deformable Objects via Physics-Aware Demonstration Synthesis
by: Yang, Yu, et al.
Published: (2025)
by: Yang, Yu, et al.
Published: (2025)
When Tracking Fails: Analyzing Failure Modes of SAM2 for Point-Based Tracking in Surgical Videos
by: Jang, Woowon, et al.
Published: (2025)
by: Jang, Woowon, et al.
Published: (2025)
OpenSDI: Spotting Diffusion-Generated Images in the Open World
by: Wang, Yabin, et al.
Published: (2025)
by: Wang, Yabin, et al.
Published: (2025)
Mamba-MOC: A Multicategory Remote Object Counting via State Space Model
by: Liu, Peng, et al.
Published: (2025)
by: Liu, Peng, et al.
Published: (2025)
EgoGrasp: World-Space Hand-Object Interaction Estimation from Egocentric Videos
by: Fu, Hongming, et al.
Published: (2026)
by: Fu, Hongming, et al.
Published: (2026)
VOID: Video Object and Interaction Deletion
by: Motamed, Saman, et al.
Published: (2026)
by: Motamed, Saman, et al.
Published: (2026)
Make It Count: Text-to-Image Generation with an Accurate Number of Objects
by: Binyamin, Lital, et al.
Published: (2024)
by: Binyamin, Lital, et al.
Published: (2024)
SOK-Bench: A Situated Video Reasoning Benchmark with Aligned Open-World Knowledge
by: Wang, Andong, et al.
Published: (2024)
by: Wang, Andong, et al.
Published: (2024)
PanoWorld: Geometry-Consistent Panoramic Video World Modeling
by: Jiang, Le, et al.
Published: (2026)
by: Jiang, Le, et al.
Published: (2026)
Temporal Object-Aware Vision Transformer for Few-Shot Video Object Detection
by: Kumar, Yogesh, et al.
Published: (2025)
by: Kumar, Yogesh, et al.
Published: (2025)
Open-Sora Plan: Open-Source Large Video Generation Model
by: Lin, Bin, et al.
Published: (2024)
by: Lin, Bin, et al.
Published: (2024)
OCTrack: Benchmarking the Open-Corpus Multi-Object Tracking
by: Qian, Zekun, et al.
Published: (2024)
by: Qian, Zekun, et al.
Published: (2024)
Sampling Bag of Views for Open-Vocabulary Object Detection
by: Choi, Hojun, et al.
Published: (2024)
by: Choi, Hojun, et al.
Published: (2024)
Similar Items
-
CountGD++: Generalized Prompting for Open-World Counting
by: Amini-Naieni, Niki, et al.
Published: (2025) -
CountGD: Multi-Modal Open-World Counting
by: Amini-Naieni, Niki, et al.
Published: (2024) -
The MixCount Dataset: Bridging the Data Gap for Open-Vocabulary Object Counting
by: Dumery, Corentin, et al.
Published: (2026) -
Unveiling Ontological Commitment in Multi-Modal Foundation Models
by: Keser, Mert, et al.
Published: (2024) -
OVR: A Dataset for Open Vocabulary Temporal Repetition Counting in Videos
by: Dwibedi, Debidatta, et al.
Published: (2024)