Object-Shot Enhanced Grounding Network for Egocentric Video

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Feng, Yisen, Zhang, Haoyu, Liu, Meng, Guan, Weili, Nie, Liqiang
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909603409690624
author Feng, Yisen
Zhang, Haoyu
Liu, Meng
Guan, Weili
Nie, Liqiang
author_facet Feng, Yisen
Zhang, Haoyu
Liu, Meng
Guan, Weili
Nie, Liqiang
contents Egocentric video grounding is a crucial task for embodied intelligence applications, distinct from exocentric video moment localization. Existing methods primarily focus on the distributional differences between egocentric and exocentric videos but often neglect key characteristics of egocentric videos and the fine-grained information emphasized by question-type queries. To address these limitations, we propose OSGNet, an Object-Shot enhanced Grounding Network for egocentric video. Specifically, we extract object information from videos to enrich video representation, particularly for objects highlighted in the textual query but not directly captured in the video features. Additionally, we analyze the frequent shot movements inherent to egocentric videos, leveraging these features to extract the wearer's attention information, which enhances the model's ability to perform modality alignment. Experiments conducted on three datasets demonstrate that OSGNet achieves state-of-the-art performance, validating the effectiveness of our approach. Our code can be found at https://github.com/Yisen-Feng/OSGNet.
format Preprint
id arxiv_https___arxiv_org_abs_2505_04270
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Object-Shot Enhanced Grounding Network for Egocentric Video
Feng, Yisen
Zhang, Haoyu
Liu, Meng
Guan, Weili
Nie, Liqiang
Computer Vision and Pattern Recognition
Artificial Intelligence
Egocentric video grounding is a crucial task for embodied intelligence applications, distinct from exocentric video moment localization. Existing methods primarily focus on the distributional differences between egocentric and exocentric videos but often neglect key characteristics of egocentric videos and the fine-grained information emphasized by question-type queries. To address these limitations, we propose OSGNet, an Object-Shot enhanced Grounding Network for egocentric video. Specifically, we extract object information from videos to enrich video representation, particularly for objects highlighted in the textual query but not directly captured in the video features. Additionally, we analyze the frequent shot movements inherent to egocentric videos, leveraging these features to extract the wearer's attention information, which enhances the model's ability to perform modality alignment. Experiments conducted on three datasets demonstrate that OSGNet achieves state-of-the-art performance, validating the effectiveness of our approach. Our code can be found at https://github.com/Yisen-Feng/OSGNet.
title Object-Shot Enhanced Grounding Network for Egocentric Video
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2505.04270