OSGNet @ Ego4D Episodic Memory Challenge 2025

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Feng, Yisen, Zhang, Haoyu, Chu, Qiaohui, Liu, Meng, Guan, Weili, Wang, Yaowei, Nie, Liqiang
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915324500115456
author Feng, Yisen
Zhang, Haoyu
Chu, Qiaohui
Liu, Meng
Guan, Weili
Wang, Yaowei
Nie, Liqiang
author_facet Feng, Yisen
Zhang, Haoyu
Chu, Qiaohui
Liu, Meng
Guan, Weili
Wang, Yaowei
Nie, Liqiang
contents In this report, we present our champion solutions for the three egocentric video localization tracks of the Ego4D Episodic Memory Challenge at CVPR 2025. All tracks require precise localization of the interval within an untrimmed egocentric video. Previous unified video localization approaches often rely on late fusion strategies, which tend to yield suboptimal results. To address this, we adopt an early fusion-based video localization model to tackle all three tasks, aiming to enhance localization accuracy. Ultimately, our method achieved first place in the Natural Language Queries, Goal Step, and Moment Queries tracks, demonstrating its effectiveness. Our code can be found at https://github.com/Yisen-Feng/OSGNet.
format Preprint
id arxiv_https___arxiv_org_abs_2506_03710
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OSGNet @ Ego4D Episodic Memory Challenge 2025
Feng, Yisen
Zhang, Haoyu
Chu, Qiaohui
Liu, Meng
Guan, Weili
Wang, Yaowei
Nie, Liqiang
Computer Vision and Pattern Recognition
Artificial Intelligence
In this report, we present our champion solutions for the three egocentric video localization tracks of the Ego4D Episodic Memory Challenge at CVPR 2025. All tracks require precise localization of the interval within an untrimmed egocentric video. Previous unified video localization approaches often rely on late fusion strategies, which tend to yield suboptimal results. To address this, we adopt an early fusion-based video localization model to tackle all three tasks, aiming to enhance localization accuracy. Ultimately, our method achieved first place in the Natural Language Queries, Goal Step, and Moment Queries tracks, demonstrating its effectiveness. Our code can be found at https://github.com/Yisen-Feng/OSGNet.
title OSGNet @ Ego4D Episodic Memory Challenge 2025
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2506.03710