Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Maaz, Muhammad, Rasheed, Hanoona, Khan, Salman, Khan, Fahad Shahbaz
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916280501534720
author Maaz, Muhammad
Rasheed, Hanoona
Khan, Salman
Khan, Fahad Shahbaz
author_facet Maaz, Muhammad
Rasheed, Hanoona
Khan, Salman
Khan, Fahad Shahbaz
contents Conversation agents fueled by Large Language Models (LLMs) are providing a new way to interact with visual data. While there have been initial attempts for image-based conversation models, this work addresses the under-explored field of \emph{video-based conversation} by introducing Video-ChatGPT. It is a multimodal model that merges a video-adapted visual encoder with an LLM. The resulting model is capable of understanding and generating detailed conversations about videos. We introduce a new dataset of 100,000 video-instruction pairs used to train Video-ChatGPT acquired via manual and semi-automated pipeline that is easily scalable and robust to label noise. We also develop a quantitative evaluation framework for video-based dialogue models to objectively analyze the strengths and weaknesses of video-based dialogue models. Code: https://github.com/mbzuai-oryx/Video-ChatGPT.
format Preprint
id arxiv_https___arxiv_org_abs_2306_05424
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Maaz, Muhammad
Rasheed, Hanoona
Khan, Salman
Khan, Fahad Shahbaz
Computer Vision and Pattern Recognition
Conversation agents fueled by Large Language Models (LLMs) are providing a new way to interact with visual data. While there have been initial attempts for image-based conversation models, this work addresses the under-explored field of \emph{video-based conversation} by introducing Video-ChatGPT. It is a multimodal model that merges a video-adapted visual encoder with an LLM. The resulting model is capable of understanding and generating detailed conversations about videos. We introduce a new dataset of 100,000 video-instruction pairs used to train Video-ChatGPT acquired via manual and semi-automated pipeline that is easily scalable and robust to label noise. We also develop a quantitative evaluation framework for video-based dialogue models to objectively analyze the strengths and weaknesses of video-based dialogue models. Code: https://github.com/mbzuai-oryx/Video-ChatGPT.
title Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2306.05424