F$^3$Set: Towards Analyzing Fast, Frequent, and Fine-grained Events from Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Zhaoyu, Jiang, Kan, Ma, Murong, Hou, Zhe, Lin, Yun, Dong, Jin Song
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917985648640000
author Liu, Zhaoyu
Jiang, Kan
Ma, Murong
Hou, Zhe
Lin, Yun
Dong, Jin Song
author_facet Liu, Zhaoyu
Jiang, Kan
Ma, Murong
Hou, Zhe
Lin, Yun
Dong, Jin Song
contents Analyzing Fast, Frequent, and Fine-grained (F$^3$) events presents a significant challenge in video analytics and multi-modal LLMs. Current methods struggle to identify events that satisfy all the F$^3$ criteria with high accuracy due to challenges such as motion blur and subtle visual discrepancies. To advance research in video understanding, we introduce F$^3$Set, a benchmark that consists of video datasets for precise F$^3$ event detection. Datasets in F$^3$Set are characterized by their extensive scale and comprehensive detail, usually encompassing over 1,000 event types with precise timestamps and supporting multi-level granularity. Currently, F$^3$Set contains several sports datasets, and this framework may be extended to other applications as well. We evaluated popular temporal action understanding methods on F$^3$Set, revealing substantial challenges for existing techniques. Additionally, we propose a new method, F$^3$ED, for F$^3$ event detections, achieving superior performance. The dataset, model, and benchmark code are available at https://github.com/F3Set/F3Set.
format Preprint
id arxiv_https___arxiv_org_abs_2504_08222
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle F$^3$Set: Towards Analyzing Fast, Frequent, and Fine-grained Events from Videos
Liu, Zhaoyu
Jiang, Kan
Ma, Murong
Hou, Zhe
Lin, Yun
Dong, Jin Song
Computer Vision and Pattern Recognition
Artificial Intelligence
Analyzing Fast, Frequent, and Fine-grained (F$^3$) events presents a significant challenge in video analytics and multi-modal LLMs. Current methods struggle to identify events that satisfy all the F$^3$ criteria with high accuracy due to challenges such as motion blur and subtle visual discrepancies. To advance research in video understanding, we introduce F$^3$Set, a benchmark that consists of video datasets for precise F$^3$ event detection. Datasets in F$^3$Set are characterized by their extensive scale and comprehensive detail, usually encompassing over 1,000 event types with precise timestamps and supporting multi-level granularity. Currently, F$^3$Set contains several sports datasets, and this framework may be extended to other applications as well. We evaluated popular temporal action understanding methods on F$^3$Set, revealing substantial challenges for existing techniques. Additionally, we propose a new method, F$^3$ED, for F$^3$ event detections, achieving superior performance. The dataset, model, and benchmark code are available at https://github.com/F3Set/F3Set.
title F$^3$Set: Towards Analyzing Fast, Frequent, and Fine-grained Events from Videos
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2504.08222