NAVERO: Unlocking Fine-Grained Semantics for Video-Language Compositionality

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Tao, Chaofan, Kwon, Gukyeong, Gunjal, Varad, Yang, Hao, Cai, Zhaowei, Dukler, Yonatan, Swaminathan, Ashwin, Manmatha, R., Taylor, Colin Jon, Soatto, Stefano
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913471121063936
author Tao, Chaofan
Kwon, Gukyeong
Gunjal, Varad
Yang, Hao
Cai, Zhaowei
Dukler, Yonatan
Swaminathan, Ashwin
Manmatha, R.
Taylor, Colin Jon
Soatto, Stefano
author_facet Tao, Chaofan
Kwon, Gukyeong
Gunjal, Varad
Yang, Hao
Cai, Zhaowei
Dukler, Yonatan
Swaminathan, Ashwin
Manmatha, R.
Taylor, Colin Jon
Soatto, Stefano
contents We study the capability of Video-Language (VidL) models in understanding compositions between objects, attributes, actions and their relations. Composition understanding becomes particularly challenging for video data since the compositional relations rapidly change over time in videos. We first build a benchmark named AARO to evaluate composition understanding related to actions on top of spatial concepts. The benchmark is constructed by generating negative texts with incorrect action descriptions for a given video and the model is expected to pair a positive text with its corresponding video. Furthermore, we propose a training method called NAVERO which utilizes video-text data augmented with negative texts to enhance composition understanding. We also develop a negative-augmented visual-language matching loss which is used explicitly to benefit from the generated negative text. We compare NAVERO with other state-of-the-art methods in terms of compositional understanding as well as video-text retrieval performance. NAVERO achieves significant improvement over other methods for both video-language and image-language composition understanding, while maintaining strong performance on traditional text-video retrieval tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2408_09511
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle NAVERO: Unlocking Fine-Grained Semantics for Video-Language Compositionality
Tao, Chaofan
Kwon, Gukyeong
Gunjal, Varad
Yang, Hao
Cai, Zhaowei
Dukler, Yonatan
Swaminathan, Ashwin
Manmatha, R.
Taylor, Colin Jon
Soatto, Stefano
Computer Vision and Pattern Recognition
We study the capability of Video-Language (VidL) models in understanding compositions between objects, attributes, actions and their relations. Composition understanding becomes particularly challenging for video data since the compositional relations rapidly change over time in videos. We first build a benchmark named AARO to evaluate composition understanding related to actions on top of spatial concepts. The benchmark is constructed by generating negative texts with incorrect action descriptions for a given video and the model is expected to pair a positive text with its corresponding video. Furthermore, we propose a training method called NAVERO which utilizes video-text data augmented with negative texts to enhance composition understanding. We also develop a negative-augmented visual-language matching loss which is used explicitly to benefit from the generated negative text. We compare NAVERO with other state-of-the-art methods in terms of compositional understanding as well as video-text retrieval performance. NAVERO achieves significant improvement over other methods for both video-language and image-language composition understanding, while maintaining strong performance on traditional text-video retrieval tasks.
title NAVERO: Unlocking Fine-Grained Semantics for Video-Language Compositionality
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2408.09511