U-CESE: Unified Clip-based Event Search Engine for AI Challenge HCMC 2025

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Le, Duc-Nhuan, Nguyen, Hoang-Phuc, Lam, Thanh-Duy, Dang, Minh-Nhut, Le, Minh-Hoang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910247567753216
author Le, Duc-Nhuan
Nguyen, Hoang-Phuc
Lam, Thanh-Duy
Dang, Minh-Nhut
Le, Minh-Hoang
author_facet Le, Duc-Nhuan
Nguyen, Hoang-Phuc
Lam, Thanh-Duy
Dang, Minh-Nhut
Le, Minh-Hoang
contents Retrieving events from large-scale video datasets is challenging due to complex temporal, spatial, and multimodal information. This paper presents U-CESE, our solution for the AI Challenge HCMC 2025, a Unified Clip-based Event Search Engine for multimodal event retrieval across diverse video sources. Building on CESE, U-CESE integrates its three modules into a single cohesive framework, ensuring consistent processing and retrieval across query types. A core component is the Unified Clipping Algorithm, which merges separate clipping algorithms into one efficient pipeline. To handle large-scale data, we propose DAKE, a lightweight, training-free keyframe extraction method using JPEG file size variations to identify significant scene changes. Finally, we introduce ReCap, a temporally consistent captioning framework inspired by Recurrent Neural Network, generating detailed and context-aware textual descriptions. Experiments show that U-CESE delivers robust, consistent, and efficient performance in large-scale multimodal event retrieval.
format Preprint
id arxiv_https___arxiv_org_abs_2605_23274
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle U-CESE: Unified Clip-based Event Search Engine for AI Challenge HCMC 2025
Le, Duc-Nhuan
Nguyen, Hoang-Phuc
Lam, Thanh-Duy
Dang, Minh-Nhut
Le, Minh-Hoang
Computer Vision and Pattern Recognition
Retrieving events from large-scale video datasets is challenging due to complex temporal, spatial, and multimodal information. This paper presents U-CESE, our solution for the AI Challenge HCMC 2025, a Unified Clip-based Event Search Engine for multimodal event retrieval across diverse video sources. Building on CESE, U-CESE integrates its three modules into a single cohesive framework, ensuring consistent processing and retrieval across query types. A core component is the Unified Clipping Algorithm, which merges separate clipping algorithms into one efficient pipeline. To handle large-scale data, we propose DAKE, a lightweight, training-free keyframe extraction method using JPEG file size variations to identify significant scene changes. Finally, we introduce ReCap, a temporally consistent captioning framework inspired by Recurrent Neural Network, generating detailed and context-aware textual descriptions. Experiments show that U-CESE delivers robust, consistent, and efficient performance in large-scale multimodal event retrieval.
title U-CESE: Unified Clip-based Event Search Engine for AI Challenge HCMC 2025
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.23274