Seeking Universal Shot Language Understanding Solutions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Haoxin, Kamarthi, Harshavardhan, Zhao, Zhiyuan, Chen, Hongjie, Prakash, B. Aditya
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908900524032000
author Liu, Haoxin
Kamarthi, Harshavardhan
Zhao, Zhiyuan
Chen, Hongjie
Prakash, B. Aditya
author_facet Liu, Haoxin
Kamarthi, Harshavardhan
Zhao, Zhiyuan
Chen, Hongjie
Prakash, B. Aditya
contents Shot language understanding (SLU) is crucial for cinematic analysis but remains challenging due to its diverse cinematographic dimensions and subjective expert judgment. While vision-language models (VLMs) have shown strong ability in general visual understanding, recent studies reveal judgment discrepancies between VLMs and film experts on SLU tasks. To address this gap, we introduce SLU-SUITE, a comprehensive training and evaluation suite containing 490K human-annotated QA pairs across 33 tasks spanning six film-grounded dimensions. Using SLU-SUITE, we originally observe two insights into VLM-based SLU from: the model side, which diagnoses key bottlenecks of modules; the data side, which quantifies cross-dimensional influences among tasks. These findings motivate our universal SLU solutions from two complementary paradigms: UniShot, a balanced one-for-all generalist trained via dynamic-balanced data mixing, and AgentShots, a prompt-routed expert cluster that maximizes peak dimension performance. Extensive experiments show that our models outperform task-specific ensembles on in-domain tasks and surpass leading commercial VLMs by 22% on out-of-domain tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2603_18448
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Seeking Universal Shot Language Understanding Solutions
Liu, Haoxin
Kamarthi, Harshavardhan
Zhao, Zhiyuan
Chen, Hongjie
Prakash, B. Aditya
Machine Learning
Shot language understanding (SLU) is crucial for cinematic analysis but remains challenging due to its diverse cinematographic dimensions and subjective expert judgment. While vision-language models (VLMs) have shown strong ability in general visual understanding, recent studies reveal judgment discrepancies between VLMs and film experts on SLU tasks. To address this gap, we introduce SLU-SUITE, a comprehensive training and evaluation suite containing 490K human-annotated QA pairs across 33 tasks spanning six film-grounded dimensions. Using SLU-SUITE, we originally observe two insights into VLM-based SLU from: the model side, which diagnoses key bottlenecks of modules; the data side, which quantifies cross-dimensional influences among tasks. These findings motivate our universal SLU solutions from two complementary paradigms: UniShot, a balanced one-for-all generalist trained via dynamic-balanced data mixing, and AgentShots, a prompt-routed expert cluster that maximizes peak dimension performance. Extensive experiments show that our models outperform task-specific ensembles on in-domain tasks and surpass leading commercial VLMs by 22% on out-of-domain tasks.
title Seeking Universal Shot Language Understanding Solutions
topic Machine Learning
url https://arxiv.org/abs/2603.18448