F$^3$Set: Towards Analyzing Fast, Frequent, and Fine-grained Events from Videos
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917985648640000 |
|---|---|
| author | Liu, Zhaoyu Jiang, Kan Ma, Murong Hou, Zhe Lin, Yun Dong, Jin Song |
| author_facet | Liu, Zhaoyu Jiang, Kan Ma, Murong Hou, Zhe Lin, Yun Dong, Jin Song |
| contents | Analyzing Fast, Frequent, and Fine-grained (F$^3$) events presents a significant challenge in video analytics and multi-modal LLMs. Current methods struggle to identify events that satisfy all the F$^3$ criteria with high accuracy due to challenges such as motion blur and subtle visual discrepancies. To advance research in video understanding, we introduce F$^3$Set, a benchmark that consists of video datasets for precise F$^3$ event detection. Datasets in F$^3$Set are characterized by their extensive scale and comprehensive detail, usually encompassing over 1,000 event types with precise timestamps and supporting multi-level granularity. Currently, F$^3$Set contains several sports datasets, and this framework may be extended to other applications as well. We evaluated popular temporal action understanding methods on F$^3$Set, revealing substantial challenges for existing techniques. Additionally, we propose a new method, F$^3$ED, for F$^3$ event detections, achieving superior performance. The dataset, model, and benchmark code are available at https://github.com/F3Set/F3Set. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_08222 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | F$^3$Set: Towards Analyzing Fast, Frequent, and Fine-grained Events from Videos Liu, Zhaoyu Jiang, Kan Ma, Murong Hou, Zhe Lin, Yun Dong, Jin Song Computer Vision and Pattern Recognition Artificial Intelligence Analyzing Fast, Frequent, and Fine-grained (F$^3$) events presents a significant challenge in video analytics and multi-modal LLMs. Current methods struggle to identify events that satisfy all the F$^3$ criteria with high accuracy due to challenges such as motion blur and subtle visual discrepancies. To advance research in video understanding, we introduce F$^3$Set, a benchmark that consists of video datasets for precise F$^3$ event detection. Datasets in F$^3$Set are characterized by their extensive scale and comprehensive detail, usually encompassing over 1,000 event types with precise timestamps and supporting multi-level granularity. Currently, F$^3$Set contains several sports datasets, and this framework may be extended to other applications as well. We evaluated popular temporal action understanding methods on F$^3$Set, revealing substantial challenges for existing techniques. Additionally, we propose a new method, F$^3$ED, for F$^3$ event detections, achieving superior performance. The dataset, model, and benchmark code are available at https://github.com/F3Set/F3Set. |
| title | F$^3$Set: Towards Analyzing Fast, Frequent, and Fine-grained Events from Videos |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2504.08222 |