Interpretable Embedding for Ad-hoc Video Search

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wu, Jiaxin, Ngo, Chong-Wah
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913237055832064
author Wu, Jiaxin
Ngo, Chong-Wah
author_facet Wu, Jiaxin
Ngo, Chong-Wah
contents Answering query with semantic concepts has long been the mainstream approach for video search. Until recently, its performance is surpassed by concept-free approach, which embeds queries in a joint space as videos. Nevertheless, the embedded features as well as search results are not interpretable, hindering subsequent steps in video browsing and query reformulation. This paper integrates feature embedding and concept interpretation into a neural network for unified dual-task learning. In this way, an embedding is associated with a list of semantic concepts as an interpretation of video content. This paper empirically demonstrates that, by using either the embedding features or concepts, considerable search improvement is attainable on TRECVid benchmarked datasets. Concepts are not only effective in pruning false positive videos, but also highly complementary to concept-free search, leading to large margin of improvement compared to state-of-the-art approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2402_11812
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Interpretable Embedding for Ad-hoc Video Search
Wu, Jiaxin
Ngo, Chong-Wah
Computer Vision and Pattern Recognition
Multimedia
Answering query with semantic concepts has long been the mainstream approach for video search. Until recently, its performance is surpassed by concept-free approach, which embeds queries in a joint space as videos. Nevertheless, the embedded features as well as search results are not interpretable, hindering subsequent steps in video browsing and query reformulation. This paper integrates feature embedding and concept interpretation into a neural network for unified dual-task learning. In this way, an embedding is associated with a list of semantic concepts as an interpretation of video content. This paper empirically demonstrates that, by using either the embedding features or concepts, considerable search improvement is attainable on TRECVid benchmarked datasets. Concepts are not only effective in pruning false positive videos, but also highly complementary to concept-free search, leading to large margin of improvement compared to state-of-the-art approaches.
title Interpretable Embedding for Ad-hoc Video Search
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2402.11812