A Short Note on Evaluating RepNet for Temporal Repetition Counting in Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Dwibedi, Debidatta, Aytar, Yusuf, Tompson, Jonathan, Sermanet, Pierre, Zisserman, Andrew |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OVR: A Dataset for Open Vocabulary Temporal Repetition Counting in Videos
by: Dwibedi, Debidatta, et al.
Published: (2024)
by: Dwibedi, Debidatta, et al.
Published: (2024)
FlexCap: Describe Anything in Images in Controllable Detail
by: Dwibedi, Debidatta, et al.
Published: (2024)
by: Dwibedi, Debidatta, et al.
Published: (2024)
Open-World Object Counting in Videos
by: Amini-Naieni, Niki, et al.
Published: (2025)
by: Amini-Naieni, Niki, et al.
Published: (2025)
RepNet-VSR: Reparameterizable Architecture for High-Fidelity Video Super-Resolution
by: Wu, Biao, et al.
Published: (2025)
by: Wu, Biao, et al.
Published: (2025)
Learning from One Continuous Video Stream
by: Carreira, João, et al.
Published: (2023)
by: Carreira, João, et al.
Published: (2023)
RT-H: Action Hierarchies Using Language
by: Belkhale, Suneel, et al.
Published: (2024)
by: Belkhale, Suneel, et al.
Published: (2024)
Geo-RepNet: Geometry-Aware Representation Learning for Surgical Phase Recognition in Endoscopic Submucosal Dissection
by: Tang, Rui, et al.
Published: (2025)
by: Tang, Rui, et al.
Published: (2025)
IVAC-P2L: Leveraging Irregular Repetition Priors for Improving Video Action Counting
by: Wang, Hang, et al.
Published: (2024)
by: Wang, Hang, et al.
Published: (2024)
CountFormer: A Transformer Framework for Learning Visual Repetition and Structure in Class-Agnostic Object Counting
by: Hossain, Md Tanvir, et al.
Published: (2025)
by: Hossain, Md Tanvir, et al.
Published: (2025)
New keypoint-based approach for recognising British Sign Language (BSL) from sequences
by: Deb, Oishi, et al.
Published: (2024)
by: Deb, Oishi, et al.
Published: (2024)
B-Rep Distance Functions (BR-DF): How to Represent a B-Rep Model by Volumetric Distance Functions?
by: Zhang, Fuyang, et al.
Published: (2025)
by: Zhang, Fuyang, et al.
Published: (2025)
TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability
by: Chen, Shimin, et al.
Published: (2024)
by: Chen, Shimin, et al.
Published: (2024)
DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long Videos
by: Lu, Zijia, et al.
Published: (2025)
by: Lu, Zijia, et al.
Published: (2025)
CountGD++: Generalized Prompting for Open-World Counting
by: Amini-Naieni, Niki, et al.
Published: (2025)
by: Amini-Naieni, Niki, et al.
Published: (2025)
3D-Aware Instance Segmentation and Tracking in Egocentric Videos
by: Bhalgat, Yash, et al.
Published: (2024)
by: Bhalgat, Yash, et al.
Published: (2024)
UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition
by: Ding, Xiaohan, et al.
Published: (2023)
by: Ding, Xiaohan, et al.
Published: (2023)
Demystifying KAN for Vision Tasks: The RepKAN Approach
by: Cheon, Minjong
Published: (2026)
by: Cheon, Minjong
Published: (2026)
RepVGG-GELAN: Enhanced GELAN with VGG-STYLE ConvNets for Brain Tumour Detection
by: Balakrishnan, Thennarasi, et al.
Published: (2024)
by: Balakrishnan, Thennarasi, et al.
Published: (2024)
Self-supervised Learning on Camera Trap Footage Yields a Strong Universal Face Embedder
by: Iashin, Vladimir, et al.
Published: (2025)
by: Iashin, Vladimir, et al.
Published: (2025)
Benefits of Feature Extraction and Temporal Sequence Analysis for Video Frame Prediction: An Evaluation of Hybrid Deep Learning Models
by: Velázquez, Jose M. Sánchez, et al.
Published: (2025)
by: Velázquez, Jose M. Sánchez, et al.
Published: (2025)
Temporally-Aware Supervised Contrastive Learning for Polyp Counting in Colonoscopy
by: Parolari, Luca, et al.
Published: (2025)
by: Parolari, Luca, et al.
Published: (2025)
SciVid: Cross-Domain Evaluation of Video Models in Scientific Applications
by: Hasson, Yana, et al.
Published: (2025)
by: Hasson, Yana, et al.
Published: (2025)
CountGD: Multi-Modal Open-World Counting
by: Amini-Naieni, Niki, et al.
Published: (2024)
by: Amini-Naieni, Niki, et al.
Published: (2024)
Can You Count to Nine? A Human Evaluation Benchmark for Counting Limits in Modern Text-to-Video Models
by: Guo, Xuyang, et al.
Published: (2025)
by: Guo, Xuyang, et al.
Published: (2025)
TempGlitch: Evaluating Vision-Language Models for Temporal Glitch Detection in Gameplay Videos
by: Yu, Yakun, et al.
Published: (2026)
by: Yu, Yakun, et al.
Published: (2026)
LOLGORITHM: Funny Comment Generation Agent For Short Videos
by: Ouyang, Xuan, et al.
Published: (2026)
by: Ouyang, Xuan, et al.
Published: (2026)
Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos
by: Fei, Jiajun, et al.
Published: (2024)
by: Fei, Jiajun, et al.
Published: (2024)
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
by: Wang, Shihao, et al.
Published: (2025)
by: Wang, Shihao, et al.
Published: (2025)
ViscoNet: Bridging and Harmonizing Visual and Textual Conditioning for ControlNet
by: Cheong, Soon Yau, et al.
Published: (2023)
by: Cheong, Soon Yau, et al.
Published: (2023)
Generating Robot Constitutions & Benchmarks for Semantic Safety
by: Sermanet, Pierre, et al.
Published: (2025)
by: Sermanet, Pierre, et al.
Published: (2025)
USV: Towards Understanding the User-generated Short-form Videos
by: Cheng, Haoyue, et al.
Published: (2026)
by: Cheng, Haoyue, et al.
Published: (2026)
RepAct: The Re-parameterizable Adaptive Activation Function
by: Wu, Xian, et al.
Published: (2024)
by: Wu, Xian, et al.
Published: (2024)
PackForcing: Short Video Training Suffices for Long Video Sampling and Long Context Inference
by: Mao, Xiaofeng, et al.
Published: (2026)
by: Mao, Xiaofeng, et al.
Published: (2026)
Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning
by: Yue, Feng, et al.
Published: (2025)
by: Yue, Feng, et al.
Published: (2025)
VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning
by: Ding, Yang, et al.
Published: (2025)
by: Ding, Yang, et al.
Published: (2025)
OccluNet: Spatio-Temporal Deep Learning for Occlusion Detection on DSA
by: Kore, Anushka A., et al.
Published: (2025)
by: Kore, Anushka A., et al.
Published: (2025)
AgRegNet: A Deep Regression Network for Flower and Fruit Density Estimation, Localization, and Counting in Orchards
by: Bhattarai, Uddhav, et al.
Published: (2024)
by: Bhattarai, Uddhav, et al.
Published: (2024)
VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos
by: Liu, Wenqi, et al.
Published: (2026)
by: Liu, Wenqi, et al.
Published: (2026)
Evaluating Gemini Robotics Policies in a Veo World Simulator
by: Gemini Robotics Team, et al.
Published: (2025)
by: Gemini Robotics Team, et al.
Published: (2025)
REGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video Editing
by: Xu, Weihan, et al.
Published: (2025)
by: Xu, Weihan, et al.
Published: (2025)
Similar Items
-
OVR: A Dataset for Open Vocabulary Temporal Repetition Counting in Videos
by: Dwibedi, Debidatta, et al.
Published: (2024) -
FlexCap: Describe Anything in Images in Controllable Detail
by: Dwibedi, Debidatta, et al.
Published: (2024) -
Open-World Object Counting in Videos
by: Amini-Naieni, Niki, et al.
Published: (2025) -
RepNet-VSR: Reparameterizable Architecture for High-Fidelity Video Super-Resolution
by: Wu, Biao, et al.
Published: (2025) -
Learning from One Continuous Video Stream
by: Carreira, João, et al.
Published: (2023)