ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Liu, Hongbo, He, Jingwen, Jin, Yi, Zheng, Dian, Dong, Yuhao, Zhang, Fan, Huang, Ziqi, He, Yinan, Li, Yangguang, Chen, Weichao, Qiao, Yu, Ouyang, Wanli, Zhao, Shengjie, Liu, Ziwei
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908918143254528
author Liu, Hongbo
He, Jingwen
Jin, Yi
Zheng, Dian
Dong, Yuhao
Zhang, Fan
Huang, Ziqi
He, Yinan
Li, Yangguang
Chen, Weichao
Qiao, Yu
Ouyang, Wanli
Zhao, Shengjie
Liu, Ziwei
author_facet Liu, Hongbo
He, Jingwen
Jin, Yi
Zheng, Dian
Dong, Yuhao
Zhang, Fan
Huang, Ziqi
He, Yinan
Li, Yangguang
Chen, Weichao
Qiao, Yu
Ouyang, Wanli
Zhao, Shengjie
Liu, Ziwei
contents Cinematography, the fundamental visual language of film, is essential for conveying narrative, emotion, and aesthetic quality. While recent Vision-Language Models (VLMs) demonstrate strong general visual understanding, their proficiency in comprehending the nuanced cinematic grammar embedded within individual shots remains largely unexplored and lacks robust evaluation. This critical gap limits both fine-grained visual comprehension and the precision of AI-assisted video generation. To address this, we introduce ShotBench, a comprehensive benchmark specifically designed for cinematic language understanding. It features over 3.5k expert-annotated QA pairs from images and video clips, meticulously curated from over 200 acclaimed (predominantly Oscar-nominated) films and spanning eight key cinematography dimensions. Our evaluation of 24 leading VLMs on ShotBench reveals their substantial limitations: even the top-performing model achieves less than 60% average accuracy, particularly struggling with fine-grained visual cues and complex spatial reasoning. To catalyze advancement in this domain, we construct ShotQA, a large-scale multimodal dataset comprising approximately 70k cinematic QA pairs. Leveraging ShotQA, we develop ShotVL through supervised fine-tuning and Group Relative Policy Optimization. ShotVL significantly outperforms all existing open-source and proprietary models on ShotBench, establishing new state-of-the-art performance. We open-source our models, data, and code to foster rapid progress in this crucial area of AI-driven cinematic understanding and generation.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21356
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models
Liu, Hongbo
He, Jingwen
Jin, Yi
Zheng, Dian
Dong, Yuhao
Zhang, Fan
Huang, Ziqi
He, Yinan
Li, Yangguang
Chen, Weichao
Qiao, Yu
Ouyang, Wanli
Zhao, Shengjie
Liu, Ziwei
Computer Vision and Pattern Recognition
Cinematography, the fundamental visual language of film, is essential for conveying narrative, emotion, and aesthetic quality. While recent Vision-Language Models (VLMs) demonstrate strong general visual understanding, their proficiency in comprehending the nuanced cinematic grammar embedded within individual shots remains largely unexplored and lacks robust evaluation. This critical gap limits both fine-grained visual comprehension and the precision of AI-assisted video generation. To address this, we introduce ShotBench, a comprehensive benchmark specifically designed for cinematic language understanding. It features over 3.5k expert-annotated QA pairs from images and video clips, meticulously curated from over 200 acclaimed (predominantly Oscar-nominated) films and spanning eight key cinematography dimensions. Our evaluation of 24 leading VLMs on ShotBench reveals their substantial limitations: even the top-performing model achieves less than 60% average accuracy, particularly struggling with fine-grained visual cues and complex spatial reasoning. To catalyze advancement in this domain, we construct ShotQA, a large-scale multimodal dataset comprising approximately 70k cinematic QA pairs. Leveraging ShotQA, we develop ShotVL through supervised fine-tuning and Group Relative Policy Optimization. ShotVL significantly outperforms all existing open-source and proprietary models on ShotBench, establishing new state-of-the-art performance. We open-source our models, data, and code to foster rapid progress in this crucial area of AI-driven cinematic understanding and generation.
title ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.21356