OVR: A Dataset for Open Vocabulary Temporal Repetition Counting in Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Dwibedi, Debidatta, Aytar, Yusuf, Tompson, Jonathan, Zisserman, Andrew |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Short Note on Evaluating RepNet for Temporal Repetition Counting in Videos
by: Dwibedi, Debidatta, et al.
Published: (2024)
by: Dwibedi, Debidatta, et al.
Published: (2024)
FlexCap: Describe Anything in Images in Controllable Detail
by: Dwibedi, Debidatta, et al.
Published: (2024)
by: Dwibedi, Debidatta, et al.
Published: (2024)
Open-World Object Counting in Videos
by: Amini-Naieni, Niki, et al.
Published: (2025)
by: Amini-Naieni, Niki, et al.
Published: (2025)
3D-Aware Instance Segmentation and Tracking in Egocentric Videos
by: Bhalgat, Yash, et al.
Published: (2024)
by: Bhalgat, Yash, et al.
Published: (2024)
Self-supervised Learning on Camera Trap Footage Yields a Strong Universal Face Embedder
by: Iashin, Vladimir, et al.
Published: (2025)
by: Iashin, Vladimir, et al.
Published: (2025)
Open Vocabulary Compositional Explanations for Neuron Alignment
by: La Rosa, Biagio, et al.
Published: (2025)
by: La Rosa, Biagio, et al.
Published: (2025)
The MixCount Dataset: Bridging the Data Gap for Open-Vocabulary Object Counting
by: Dumery, Corentin, et al.
Published: (2026)
by: Dumery, Corentin, et al.
Published: (2026)
LVD-2M: A Long-take Video Dataset with Temporally Dense Captions
by: Xiong, Tianwei, et al.
Published: (2024)
by: Xiong, Tianwei, et al.
Published: (2024)
Learning from One Continuous Video Stream
by: Carreira, João, et al.
Published: (2023)
by: Carreira, João, et al.
Published: (2023)
Region-centric Image-Language Pretraining for Open-Vocabulary Detection
by: Kim, Dahun, et al.
Published: (2023)
by: Kim, Dahun, et al.
Published: (2023)
EgoSurgery-HTS: A Dataset for Egocentric Hand-Tool Segmentation in Open Surgery Videos
by: Darjana, Nathan, et al.
Published: (2025)
by: Darjana, Nathan, et al.
Published: (2025)
Neural Assets: 3D-Aware Multi-Object Scene Synthesis with Image Diffusion Models
by: Wu, Ziyi, et al.
Published: (2024)
by: Wu, Ziyi, et al.
Published: (2024)
EgoSurgery-Phase: A Dataset of Surgical Phase Recognition from Egocentric Open Surgery Videos
by: Fujii, Ryo, et al.
Published: (2024)
by: Fujii, Ryo, et al.
Published: (2024)
OpenSUN3D: 1st Workshop Challenge on Open-Vocabulary 3D Scene Understanding
by: Engelmann, Francis, et al.
Published: (2024)
by: Engelmann, Francis, et al.
Published: (2024)
Can You Count to Nine? A Human Evaluation Benchmark for Counting Limits in Modern Text-to-Video Models
by: Guo, Xuyang, et al.
Published: (2025)
by: Guo, Xuyang, et al.
Published: (2025)
EgoSurgery-Tool: A Dataset of Surgical Tool and Hand Detection from Egocentric Open Surgery Videos
by: Fujii, Ryo, et al.
Published: (2024)
by: Fujii, Ryo, et al.
Published: (2024)
OV-Stitcher: A Global Context-Aware Framework for Training-Free Open-Vocabulary Semantic Segmentation
by: Moon, Seungjae, et al.
Published: (2026)
by: Moon, Seungjae, et al.
Published: (2026)
TAPVid-3D: A Benchmark for Tracking Any Point in 3D
by: Koppula, Skanda, et al.
Published: (2024)
by: Koppula, Skanda, et al.
Published: (2024)
SciVid: Cross-Domain Evaluation of Video Models in Scientific Applications
by: Hasson, Yana, et al.
Published: (2025)
by: Hasson, Yana, et al.
Published: (2025)
Incremental Learning with Repetition via Pseudo-Feature Projection
by: Tscheschner, Benedikt, et al.
Published: (2025)
by: Tscheschner, Benedikt, et al.
Published: (2025)
N2F2: Hierarchical Scene Understanding with Nested Neural Feature Fields
by: Bhalgat, Yash, et al.
Published: (2024)
by: Bhalgat, Yash, et al.
Published: (2024)
Open-Vocabulary Panoptic Segmentation Using BERT Pre-Training of Vision-Language Multiway Transformer Model
by: Chen, Yi-Chia, et al.
Published: (2024)
by: Chen, Yi-Chia, et al.
Published: (2024)
Real-Time Fitness Exercise Classification and Counting from Video Frames
by: Riccio, Riccardo
Published: (2024)
by: Riccio, Riccardo
Published: (2024)
Closed-Loop Open-Vocabulary Mobile Manipulation with GPT-4V
by: Zhi, Peiyuan, et al.
Published: (2024)
by: Zhi, Peiyuan, et al.
Published: (2024)
PLUM: Improving Inference Efficiency By Leveraging Repetition-Sparsity Trade-Off
by: Kuhar, Sachit, et al.
Published: (2023)
by: Kuhar, Sachit, et al.
Published: (2023)
CVA: Context-aware Video-text Alignment for Video Temporal Grounding
by: Moon, Sungho, et al.
Published: (2026)
by: Moon, Sungho, et al.
Published: (2026)
Deep Learning for Micro-Scale Crack Detection on Imbalanced Datasets Using Key Point Localization
by: Moreh, Fatahlla, et al.
Published: (2024)
by: Moreh, Fatahlla, et al.
Published: (2024)
VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM
by: Yuan, Yuqian, et al.
Published: (2024)
by: Yuan, Yuqian, et al.
Published: (2024)
VidTok: A Versatile and Open-Source Video Tokenizer
by: Tang, Anni, et al.
Published: (2024)
by: Tang, Anni, et al.
Published: (2024)
CountCLIP -- [Re] Teaching CLIP to Count to Ten
by: Mestha, Harshvardhan, et al.
Published: (2024)
by: Mestha, Harshvardhan, et al.
Published: (2024)
Unified Spatio-Temporal Token Scoring for Efficient Video VLMs
by: Zhang, Jianrui, et al.
Published: (2026)
by: Zhang, Jianrui, et al.
Published: (2026)
TempoControl: Temporal Attention Guidance for Text-to-Video Models
by: Schiber, Shira, et al.
Published: (2025)
by: Schiber, Shira, et al.
Published: (2025)
TA-Prompting: Enhancing Video Large Language Models for Dense Video Captioning via Temporal Anchors
by: Cheng, Wei-Yuan, et al.
Published: (2026)
by: Cheng, Wei-Yuan, et al.
Published: (2026)
Video Understanding by Design: How Datasets Shape Architectures and Insights
by: Wang, Lei, et al.
Published: (2025)
by: Wang, Lei, et al.
Published: (2025)
Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video
by: Joseph, Sonia, et al.
Published: (2025)
by: Joseph, Sonia, et al.
Published: (2025)
Evaluating Gemini Robotics Policies in a Veo World Simulator
by: Gemini Robotics Team, et al.
Published: (2025)
by: Gemini Robotics Team, et al.
Published: (2025)
Temporal-consistent CAMs for Weakly Supervised Video Segmentation in Waste Sorting
by: Marelli, Andrea, et al.
Published: (2025)
by: Marelli, Andrea, et al.
Published: (2025)
Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets
by: Nezhurina, Marianna, et al.
Published: (2025)
by: Nezhurina, Marianna, et al.
Published: (2025)
LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models
by: Helff, Lukas, et al.
Published: (2024)
by: Helff, Lukas, et al.
Published: (2024)
Video Anomaly Detection via Spatio-Temporal Pseudo-Anomaly Generation : A Unified Approach
by: Rai, Ayush K., et al.
Published: (2023)
by: Rai, Ayush K., et al.
Published: (2023)
Similar Items
-
A Short Note on Evaluating RepNet for Temporal Repetition Counting in Videos
by: Dwibedi, Debidatta, et al.
Published: (2024) -
FlexCap: Describe Anything in Images in Controllable Detail
by: Dwibedi, Debidatta, et al.
Published: (2024) -
Open-World Object Counting in Videos
by: Amini-Naieni, Niki, et al.
Published: (2025) -
3D-Aware Instance Segmentation and Tracking in Egocentric Videos
by: Bhalgat, Yash, et al.
Published: (2024) -
Self-supervised Learning on Camera Trap Footage Yields a Strong Universal Face Embedder
by: Iashin, Vladimir, et al.
Published: (2025)