ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866908918143254528 |
|---|---|
| author | Liu, Hongbo He, Jingwen Jin, Yi Zheng, Dian Dong, Yuhao Zhang, Fan Huang, Ziqi He, Yinan Li, Yangguang Chen, Weichao Qiao, Yu Ouyang, Wanli Zhao, Shengjie Liu, Ziwei |
| author_facet | Liu, Hongbo He, Jingwen Jin, Yi Zheng, Dian Dong, Yuhao Zhang, Fan Huang, Ziqi He, Yinan Li, Yangguang Chen, Weichao Qiao, Yu Ouyang, Wanli Zhao, Shengjie Liu, Ziwei |
| contents | Cinematography, the fundamental visual language of film, is essential for conveying narrative, emotion, and aesthetic quality. While recent Vision-Language Models (VLMs) demonstrate strong general visual understanding, their proficiency in comprehending the nuanced cinematic grammar embedded within individual shots remains largely unexplored and lacks robust evaluation. This critical gap limits both fine-grained visual comprehension and the precision of AI-assisted video generation. To address this, we introduce ShotBench, a comprehensive benchmark specifically designed for cinematic language understanding. It features over 3.5k expert-annotated QA pairs from images and video clips, meticulously curated from over 200 acclaimed (predominantly Oscar-nominated) films and spanning eight key cinematography dimensions. Our evaluation of 24 leading VLMs on ShotBench reveals their substantial limitations: even the top-performing model achieves less than 60% average accuracy, particularly struggling with fine-grained visual cues and complex spatial reasoning. To catalyze advancement in this domain, we construct ShotQA, a large-scale multimodal dataset comprising approximately 70k cinematic QA pairs. Leveraging ShotQA, we develop ShotVL through supervised fine-tuning and Group Relative Policy Optimization. ShotVL significantly outperforms all existing open-source and proprietary models on ShotBench, establishing new state-of-the-art performance. We open-source our models, data, and code to foster rapid progress in this crucial area of AI-driven cinematic understanding and generation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_21356 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models Liu, Hongbo He, Jingwen Jin, Yi Zheng, Dian Dong, Yuhao Zhang, Fan Huang, Ziqi He, Yinan Li, Yangguang Chen, Weichao Qiao, Yu Ouyang, Wanli Zhao, Shengjie Liu, Ziwei Computer Vision and Pattern Recognition Cinematography, the fundamental visual language of film, is essential for conveying narrative, emotion, and aesthetic quality. While recent Vision-Language Models (VLMs) demonstrate strong general visual understanding, their proficiency in comprehending the nuanced cinematic grammar embedded within individual shots remains largely unexplored and lacks robust evaluation. This critical gap limits both fine-grained visual comprehension and the precision of AI-assisted video generation. To address this, we introduce ShotBench, a comprehensive benchmark specifically designed for cinematic language understanding. It features over 3.5k expert-annotated QA pairs from images and video clips, meticulously curated from over 200 acclaimed (predominantly Oscar-nominated) films and spanning eight key cinematography dimensions. Our evaluation of 24 leading VLMs on ShotBench reveals their substantial limitations: even the top-performing model achieves less than 60% average accuracy, particularly struggling with fine-grained visual cues and complex spatial reasoning. To catalyze advancement in this domain, we construct ShotQA, a large-scale multimodal dataset comprising approximately 70k cinematic QA pairs. Leveraging ShotQA, we develop ShotVL through supervised fine-tuning and Group Relative Policy Optimization. ShotVL significantly outperforms all existing open-source and proprietary models on ShotBench, establishing new state-of-the-art performance. We open-source our models, data, and code to foster rapid progress in this crucial area of AI-driven cinematic understanding and generation. |
| title | ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2506.21356 |