OSGNet with MLLM Reranking @ Ego4D Episodic Memory Challenge 2026

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Yisen, Qu, Leigang, Zhang, Haoyu, Chu, Qiaohui, Liu, Meng, Song, Xuemeng, Guan, Weili, Nie, Liqiang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916030943592448
author Feng, Yisen
Qu, Leigang
Zhang, Haoyu
Chu, Qiaohui
Liu, Meng
Song, Xuemeng
Guan, Weili
Nie, Liqiang
author_facet Feng, Yisen
Qu, Leigang
Zhang, Haoyu
Chu, Qiaohui
Liu, Meng
Song, Xuemeng
Guan, Weili
Nie, Liqiang
contents In this report, we present our champion solutions for the Natural Language Queries and GoalStep tracks of the Ego4D Episodic Memory Challenge at CVPR 2026. Both tracks require accurately localizing temporal segments from long untrimmed egocentric videos. To address these tasks, we propose a reranking-based framework that effectively leverages the strong video-language reasoning capability of multimodal large language model (MLLM) while preserving the efficiency and candidate recall of conventional localization pipelines. Specifically, we first obtain a set of candidate segments from existing localization model OSGNet, and then employ MLLM to select the segment that best matches the given query, thereby refining the final prediction. Ultimately, our method achieved first place in both the Natural Language Queries and GoalStep tracks. Our code can be found at https://github.com/iLearn-Lab/CVPR25-OSGNet.
format Preprint
id arxiv_https___arxiv_org_abs_2605_20818
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle OSGNet with MLLM Reranking @ Ego4D Episodic Memory Challenge 2026
Feng, Yisen
Qu, Leigang
Zhang, Haoyu
Chu, Qiaohui
Liu, Meng
Song, Xuemeng
Guan, Weili
Nie, Liqiang
Computer Vision and Pattern Recognition
In this report, we present our champion solutions for the Natural Language Queries and GoalStep tracks of the Ego4D Episodic Memory Challenge at CVPR 2026. Both tracks require accurately localizing temporal segments from long untrimmed egocentric videos. To address these tasks, we propose a reranking-based framework that effectively leverages the strong video-language reasoning capability of multimodal large language model (MLLM) while preserving the efficiency and candidate recall of conventional localization pipelines. Specifically, we first obtain a set of candidate segments from existing localization model OSGNet, and then employ MLLM to select the segment that best matches the given query, thereby refining the final prediction. Ultimately, our method achieved first place in both the Natural Language Queries and GoalStep tracks. Our code can be found at https://github.com/iLearn-Lab/CVPR25-OSGNet.
title OSGNet with MLLM Reranking @ Ego4D Episodic Memory Challenge 2026
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.20818