U-CESE: Unified Clip-based Event Search Engine for AI Challenge HCMC 2025
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910247567753216 |
|---|---|
| author | Le, Duc-Nhuan Nguyen, Hoang-Phuc Lam, Thanh-Duy Dang, Minh-Nhut Le, Minh-Hoang |
| author_facet | Le, Duc-Nhuan Nguyen, Hoang-Phuc Lam, Thanh-Duy Dang, Minh-Nhut Le, Minh-Hoang |
| contents | Retrieving events from large-scale video datasets is challenging due to complex temporal, spatial, and multimodal information. This paper presents U-CESE, our solution for the AI Challenge HCMC 2025, a Unified Clip-based Event Search Engine for multimodal event retrieval across diverse video sources. Building on CESE, U-CESE integrates its three modules into a single cohesive framework, ensuring consistent processing and retrieval across query types. A core component is the Unified Clipping Algorithm, which merges separate clipping algorithms into one efficient pipeline. To handle large-scale data, we propose DAKE, a lightweight, training-free keyframe extraction method using JPEG file size variations to identify significant scene changes. Finally, we introduce ReCap, a temporally consistent captioning framework inspired by Recurrent Neural Network, generating detailed and context-aware textual descriptions. Experiments show that U-CESE delivers robust, consistent, and efficient performance in large-scale multimodal event retrieval. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_23274 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | U-CESE: Unified Clip-based Event Search Engine for AI Challenge HCMC 2025 Le, Duc-Nhuan Nguyen, Hoang-Phuc Lam, Thanh-Duy Dang, Minh-Nhut Le, Minh-Hoang Computer Vision and Pattern Recognition Retrieving events from large-scale video datasets is challenging due to complex temporal, spatial, and multimodal information. This paper presents U-CESE, our solution for the AI Challenge HCMC 2025, a Unified Clip-based Event Search Engine for multimodal event retrieval across diverse video sources. Building on CESE, U-CESE integrates its three modules into a single cohesive framework, ensuring consistent processing and retrieval across query types. A core component is the Unified Clipping Algorithm, which merges separate clipping algorithms into one efficient pipeline. To handle large-scale data, we propose DAKE, a lightweight, training-free keyframe extraction method using JPEG file size variations to identify significant scene changes. Finally, we introduce ReCap, a temporally consistent captioning framework inspired by Recurrent Neural Network, generating detailed and context-aware textual descriptions. Experiments show that U-CESE delivers robust, consistent, and efficient performance in large-scale multimodal event retrieval. |
| title | U-CESE: Unified Clip-based Event Search Engine for AI Challenge HCMC 2025 |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2605.23274 |