It's Just Another Day: Unique Video Captioning by Discriminative Prompting

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Perrett, Toby, Han, Tengda, Damen, Dima, Zisserman, Andrew
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913547649286144
author Perrett, Toby
Han, Tengda
Damen, Dima
Zisserman, Andrew
author_facet Perrett, Toby
Han, Tengda
Damen, Dima
Zisserman, Andrew
contents Long videos contain many repeating actions, events and shots. These repetitions are frequently given identical captions, which makes it difficult to retrieve the exact desired clip using a text search. In this paper, we formulate the problem of unique captioning: Given multiple clips with the same caption, we generate a new caption for each clip that uniquely identifies it. We propose Captioning by Discriminative Prompting (CDP), which predicts a property that can separate identically captioned clips, and use it to generate unique captions. We introduce two benchmarks for unique captioning, based on egocentric footage and timeloop movies - where repeating actions are common. We demonstrate that captions generated by CDP improve text-to-video R@1 by 15% for egocentric videos and 10% in timeloop movies.
format Preprint
id arxiv_https___arxiv_org_abs_2410_11702
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle It's Just Another Day: Unique Video Captioning by Discriminative Prompting
Perrett, Toby
Han, Tengda
Damen, Dima
Zisserman, Andrew
Computer Vision and Pattern Recognition
Long videos contain many repeating actions, events and shots. These repetitions are frequently given identical captions, which makes it difficult to retrieve the exact desired clip using a text search. In this paper, we formulate the problem of unique captioning: Given multiple clips with the same caption, we generate a new caption for each clip that uniquely identifies it. We propose Captioning by Discriminative Prompting (CDP), which predicts a property that can separate identically captioned clips, and use it to generate unique captions. We introduce two benchmarks for unique captioning, based on egocentric footage and timeloop movies - where repeating actions are common. We demonstrate that captions generated by CDP improve text-to-video R@1 by 15% for egocentric videos and 10% in timeloop movies.
title It's Just Another Day: Unique Video Captioning by Discriminative Prompting
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.11702