CinePile: A Long Video Question Answering Dataset and Benchmark

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rawal, Ruchit, Saifullah, Khalid, Farré, Miquel, Basri, Ronen, Jacobs, David, Somepalli, Gowthami, Goldstein, Tom
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916447058395136
author Rawal, Ruchit
Saifullah, Khalid
Farré, Miquel
Basri, Ronen
Jacobs, David
Somepalli, Gowthami
Goldstein, Tom
author_facet Rawal, Ruchit
Saifullah, Khalid
Farré, Miquel
Basri, Ronen
Jacobs, David
Somepalli, Gowthami
Goldstein, Tom
contents Current datasets for long-form video understanding often fall short of providing genuine long-form comprehension challenges, as many tasks derived from these datasets can be successfully tackled by analyzing just one or a few random frames from a video. To address this issue, we present a novel dataset and benchmark, CinePile, specifically designed for authentic long-form video understanding. This paper details our innovative approach for creating a question-answer dataset, utilizing advanced LLMs with human-in-the-loop and building upon human-generated raw data. Our comprehensive dataset comprises 305,000 multiple-choice questions (MCQs), covering various visual and multimodal aspects, including temporal comprehension, understanding human-object interactions, and reasoning about events or actions within a scene. Additionally, we fine-tuned open-source Video-LLMs on the training split and evaluated both open-source and proprietary video-centric LLMs on the test split of our dataset. The findings indicate that although current models underperform compared to humans, fine-tuning these models can lead to significant improvements in their performance.
format Preprint
id arxiv_https___arxiv_org_abs_2405_08813
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CinePile: A Long Video Question Answering Dataset and Benchmark
Rawal, Ruchit
Saifullah, Khalid
Farré, Miquel
Basri, Ronen
Jacobs, David
Somepalli, Gowthami
Goldstein, Tom
Computer Vision and Pattern Recognition
Machine Learning
Multimedia
Current datasets for long-form video understanding often fall short of providing genuine long-form comprehension challenges, as many tasks derived from these datasets can be successfully tackled by analyzing just one or a few random frames from a video. To address this issue, we present a novel dataset and benchmark, CinePile, specifically designed for authentic long-form video understanding. This paper details our innovative approach for creating a question-answer dataset, utilizing advanced LLMs with human-in-the-loop and building upon human-generated raw data. Our comprehensive dataset comprises 305,000 multiple-choice questions (MCQs), covering various visual and multimodal aspects, including temporal comprehension, understanding human-object interactions, and reasoning about events or actions within a scene. Additionally, we fine-tuned open-source Video-LLMs on the training split and evaluated both open-source and proprietary video-centric LLMs on the test split of our dataset. The findings indicate that although current models underperform compared to humans, fine-tuning these models can lead to significant improvements in their performance.
title CinePile: A Long Video Question Answering Dataset and Benchmark
topic Computer Vision and Pattern Recognition
Machine Learning
Multimedia
url https://arxiv.org/abs/2405.08813