Video ReCap: Recursive Captioning of Hour-Long Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Islam, Md Mohaiminul, Ho, Ngan, Yang, Xitong, Nagarajan, Tushar, Torresani, Lorenzo, Bertasius, Gedas
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929346650832896
author Islam, Md Mohaiminul
Ho, Ngan
Yang, Xitong
Nagarajan, Tushar
Torresani, Lorenzo
Bertasius, Gedas
author_facet Islam, Md Mohaiminul
Ho, Ngan
Yang, Xitong
Nagarajan, Tushar
Torresani, Lorenzo
Bertasius, Gedas
contents Most video captioning models are designed to process short video clips of few seconds and output text describing low-level visual concepts (e.g., objects, scenes, atomic actions). However, most real-world videos last for minutes or hours and have a complex hierarchical structure spanning different temporal granularities. We propose Video ReCap, a recursive video captioning model that can process video inputs of dramatically different lengths (from 1 second to 2 hours) and output video captions at multiple hierarchy levels. The recursive video-language architecture exploits the synergy between different video hierarchies and can process hour-long videos efficiently. We utilize a curriculum learning training scheme to learn the hierarchical structure of videos, starting from clip-level captions describing atomic actions, then focusing on segment-level descriptions, and concluding with generating summaries for hour-long videos. Furthermore, we introduce Ego4D-HCap dataset by augmenting Ego4D with 8,267 manually collected long-range video summaries. Our recursive model can flexibly generate captions at different hierarchy levels while also being useful for other complex video understanding tasks, such as VideoQA on EgoSchema. Data, code, and models are available at: https://sites.google.com/view/vidrecap
format Preprint
id arxiv_https___arxiv_org_abs_2402_13250
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Video ReCap: Recursive Captioning of Hour-Long Videos
Islam, Md Mohaiminul
Ho, Ngan
Yang, Xitong
Nagarajan, Tushar
Torresani, Lorenzo
Bertasius, Gedas
Computer Vision and Pattern Recognition
Most video captioning models are designed to process short video clips of few seconds and output text describing low-level visual concepts (e.g., objects, scenes, atomic actions). However, most real-world videos last for minutes or hours and have a complex hierarchical structure spanning different temporal granularities. We propose Video ReCap, a recursive video captioning model that can process video inputs of dramatically different lengths (from 1 second to 2 hours) and output video captions at multiple hierarchy levels. The recursive video-language architecture exploits the synergy between different video hierarchies and can process hour-long videos efficiently. We utilize a curriculum learning training scheme to learn the hierarchical structure of videos, starting from clip-level captions describing atomic actions, then focusing on segment-level descriptions, and concluding with generating summaries for hour-long videos. Furthermore, we introduce Ego4D-HCap dataset by augmenting Ego4D with 8,267 manually collected long-range video summaries. Our recursive model can flexibly generate captions at different hierarchy levels while also being useful for other complex video understanding tasks, such as VideoQA on EgoSchema. Data, code, and models are available at: https://sites.google.com/view/vidrecap
title Video ReCap: Recursive Captioning of Hour-Long Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2402.13250