Video Enriched Retrieval Augmented Generation Using Aligned Video Captions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Rosa, Kevin Dela
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914814263033856
author Rosa, Kevin Dela
author_facet Rosa, Kevin Dela
contents In this work, we propose the use of "aligned visual captions" as a mechanism for integrating information contained within videos into retrieval augmented generation (RAG) based chat assistant systems. These captions are able to describe the visual and audio content of videos in a large corpus while having the advantage of being in a textual format that is both easy to reason about & incorporate into large language model (LLM) prompts, but also typically require less multimedia content to be inserted into the multimodal LLM context window, where typical configurations can aggressively fill up the context window by sampling video frames from the source video. Furthermore, visual captions can be adapted to specific use cases by prompting the original foundational model / captioner for particular visual details or fine tuning. In hopes of helping advancing progress in this area, we curate a dataset and describe automatic evaluation procedures on common RAG tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2405_17706
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Video Enriched Retrieval Augmented Generation Using Aligned Video Captions
Rosa, Kevin Dela
Artificial Intelligence
Computer Vision and Pattern Recognition
Information Retrieval
In this work, we propose the use of "aligned visual captions" as a mechanism for integrating information contained within videos into retrieval augmented generation (RAG) based chat assistant systems. These captions are able to describe the visual and audio content of videos in a large corpus while having the advantage of being in a textual format that is both easy to reason about & incorporate into large language model (LLM) prompts, but also typically require less multimedia content to be inserted into the multimodal LLM context window, where typical configurations can aggressively fill up the context window by sampling video frames from the source video. Furthermore, visual captions can be adapted to specific use cases by prompting the original foundational model / captioner for particular visual details or fine tuning. In hopes of helping advancing progress in this area, we curate a dataset and describe automatic evaluation procedures on common RAG tasks.
title Video Enriched Retrieval Augmented Generation Using Aligned Video Captions
topic Artificial Intelligence
Computer Vision and Pattern Recognition
Information Retrieval
url https://arxiv.org/abs/2405.17706