Retrieval-Augmented Egocentric Video Captioning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Jilan, Huang, Yifei, Hou, Junlin, Chen, Guo, Zhang, Yuejie, Feng, Rui, Xie, Weidi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909227514068992
author Xu, Jilan
Huang, Yifei
Hou, Junlin
Chen, Guo
Zhang, Yuejie
Feng, Rui
Xie, Weidi
author_facet Xu, Jilan
Huang, Yifei
Hou, Junlin
Chen, Guo
Zhang, Yuejie
Feng, Rui
Xie, Weidi
contents Understanding human actions from videos of first-person view poses significant challenges. Most prior approaches explore representation learning on egocentric videos only, while overlooking the potential benefit of exploiting existing large-scale third-person videos. In this paper, (1) we develop EgoInstructor, a retrieval-augmented multimodal captioning model that automatically retrieves semantically relevant third-person instructional videos to enhance the video captioning of egocentric videos. (2) For training the cross-view retrieval module, we devise an automatic pipeline to discover ego-exo video pairs from distinct large-scale egocentric and exocentric datasets. (3) We train the cross-view retrieval module with a novel EgoExoNCE loss that pulls egocentric and exocentric video features closer by aligning them to shared text features that describe similar actions. (4) Through extensive experiments, our cross-view retrieval module demonstrates superior performance across seven benchmarks. Regarding egocentric video captioning, EgoInstructor exhibits significant improvements by leveraging third-person videos as references. Project page is available at: https://jazzcharles.github.io/Egoinstructor/
format Preprint
id arxiv_https___arxiv_org_abs_2401_00789
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Retrieval-Augmented Egocentric Video Captioning
Xu, Jilan
Huang, Yifei
Hou, Junlin
Chen, Guo
Zhang, Yuejie
Feng, Rui
Xie, Weidi
Computer Vision and Pattern Recognition
Understanding human actions from videos of first-person view poses significant challenges. Most prior approaches explore representation learning on egocentric videos only, while overlooking the potential benefit of exploiting existing large-scale third-person videos. In this paper, (1) we develop EgoInstructor, a retrieval-augmented multimodal captioning model that automatically retrieves semantically relevant third-person instructional videos to enhance the video captioning of egocentric videos. (2) For training the cross-view retrieval module, we devise an automatic pipeline to discover ego-exo video pairs from distinct large-scale egocentric and exocentric datasets. (3) We train the cross-view retrieval module with a novel EgoExoNCE loss that pulls egocentric and exocentric video features closer by aligning them to shared text features that describe similar actions. (4) Through extensive experiments, our cross-view retrieval module demonstrates superior performance across seven benchmarks. Regarding egocentric video captioning, EgoInstructor exhibits significant improvements by leveraging third-person videos as references. Project page is available at: https://jazzcharles.github.io/Egoinstructor/
title Retrieval-Augmented Egocentric Video Captioning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2401.00789