Zero-shot Audio Topic Reranking using Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Qian, Mengjie, Ma, Rao, Liusie, Adian, Loweimi, Erfan, Knill, Kate M., Gales, Mark J. F.
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916387230842880
author Qian, Mengjie
Ma, Rao
Liusie, Adian
Loweimi, Erfan
Knill, Kate M.
Gales, Mark J. F.
author_facet Qian, Mengjie
Ma, Rao
Liusie, Adian
Loweimi, Erfan
Knill, Kate M.
Gales, Mark J. F.
contents Multimodal Video Search by Examples (MVSE) investigates using video clips as the query term for information retrieval, rather than the more traditional text query. This enables far richer search modalities such as images, speaker, content, topic, and emotion. A key element for this process is highly rapid and flexible search to support large archives, which in MVSE is facilitated by representing video attributes with embeddings. This work aims to compensate for any performance loss from this rapid archive search by examining reranking approaches. In particular, zero-shot reranking methods using large language models (LLMs) are investigated as these are applicable to any video archive audio content. Performance is evaluated for topic-based retrieval on a publicly available video archive, the BBC Rewind corpus. Results demonstrate that reranking significantly improves retrieval ranking without requiring any task-specific in-domain training data. Furthermore, three sources of information (ASR transcriptions, automatic summaries and synopses) as input for LLM reranking were compared. To gain a deeper understanding and further insights into the performance differences and limitations of these text sources, we employ a fact-checking approach to analyse the information consistency among them.
format Preprint
id arxiv_https___arxiv_org_abs_2309_07606
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Zero-shot Audio Topic Reranking using Large Language Models
Qian, Mengjie
Ma, Rao
Liusie, Adian
Loweimi, Erfan
Knill, Kate M.
Gales, Mark J. F.
Computation and Language
Information Retrieval
Multimodal Video Search by Examples (MVSE) investigates using video clips as the query term for information retrieval, rather than the more traditional text query. This enables far richer search modalities such as images, speaker, content, topic, and emotion. A key element for this process is highly rapid and flexible search to support large archives, which in MVSE is facilitated by representing video attributes with embeddings. This work aims to compensate for any performance loss from this rapid archive search by examining reranking approaches. In particular, zero-shot reranking methods using large language models (LLMs) are investigated as these are applicable to any video archive audio content. Performance is evaluated for topic-based retrieval on a publicly available video archive, the BBC Rewind corpus. Results demonstrate that reranking significantly improves retrieval ranking without requiring any task-specific in-domain training data. Furthermore, three sources of information (ASR transcriptions, automatic summaries and synopses) as input for LLM reranking were compared. To gain a deeper understanding and further insights into the performance differences and limitations of these text sources, we employ a fact-checking approach to analyse the information consistency among them.
title Zero-shot Audio Topic Reranking using Large Language Models
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2309.07606