RefineShot: Rethinking Cinematography Understanding with Foundational Skill Evaluation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wu, Hang, Cai, Yujun, Ge, Haonan, Chen, Hongkai, Yang, Ming-Hsuan, Wang, Yiwei
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914071827185664
author Wu, Hang
Cai, Yujun
Ge, Haonan
Chen, Hongkai
Yang, Ming-Hsuan
Wang, Yiwei
author_facet Wu, Hang
Cai, Yujun
Ge, Haonan
Chen, Hongkai
Yang, Ming-Hsuan
Wang, Yiwei
contents Cinematography understanding refers to the ability to recognize not only the visual content of a scene but also the cinematic techniques that shape narrative meaning. This capability is attracting increasing attention, as it enhances multimodal understanding in real-world applications and underpins coherent content creation in film and media. As the most comprehensive benchmark for this task, ShotBench spans a wide range of cinematic concepts and VQA-style evaluations, with ShotVL achieving state-of-the-art results on it. However, our analysis reveals that ambiguous option design in ShotBench and ShotVL's shortcomings in reasoning consistency and instruction adherence undermine evaluation reliability, limiting fair comparison and hindering future progress. To overcome these issues, we systematically refine ShotBench through consistent option restructuring, conduct the first critical analysis of ShotVL's reasoning behavior, and introduce an extended evaluation protocol that jointly assesses task accuracy and core model competencies. These efforts lead to RefineShot, a refined and expanded benchmark that enables more reliable assessment and fosters future advances in cinematography understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2510_02423
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RefineShot: Rethinking Cinematography Understanding with Foundational Skill Evaluation
Wu, Hang
Cai, Yujun
Ge, Haonan
Chen, Hongkai
Yang, Ming-Hsuan
Wang, Yiwei
Artificial Intelligence
Cinematography understanding refers to the ability to recognize not only the visual content of a scene but also the cinematic techniques that shape narrative meaning. This capability is attracting increasing attention, as it enhances multimodal understanding in real-world applications and underpins coherent content creation in film and media. As the most comprehensive benchmark for this task, ShotBench spans a wide range of cinematic concepts and VQA-style evaluations, with ShotVL achieving state-of-the-art results on it. However, our analysis reveals that ambiguous option design in ShotBench and ShotVL's shortcomings in reasoning consistency and instruction adherence undermine evaluation reliability, limiting fair comparison and hindering future progress. To overcome these issues, we systematically refine ShotBench through consistent option restructuring, conduct the first critical analysis of ShotVL's reasoning behavior, and introduce an extended evaluation protocol that jointly assesses task accuracy and core model competencies. These efforts lead to RefineShot, a refined and expanded benchmark that enables more reliable assessment and fosters future advances in cinematography understanding.
title RefineShot: Rethinking Cinematography Understanding with Foundational Skill Evaluation
topic Artificial Intelligence
url https://arxiv.org/abs/2510.02423