Cinéaste: A Fine-grained Contextual Movie Question Answering Benchmark
Fuente:
arXiv
Salvato in:
| Autori principali: | Shah, Nisarg A., Ziai, Amir, Ekanadham, Chaitanya, Patel, Vishal M. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VidNum-1.4K: A Comprehensive Benchmark for Video-based Numerical Reasoning
di: Cui, Shaoyang, et al.
Pubblicazione: (2026)
di: Cui, Shaoyang, et al.
Pubblicazione: (2026)
StoryMovie: A Dataset for Semantic Alignment of Visual Stories with Movie Scripts and Subtitles
di: Oliveira, Daniel, et al.
Pubblicazione: (2026)
di: Oliveira, Daniel, et al.
Pubblicazione: (2026)
ReSpace: Text-Driven Autoregressive 3D Indoor Scene Synthesis and Editing
di: Bucher, Martin JJ., et al.
Pubblicazione: (2025)
di: Bucher, Martin JJ., et al.
Pubblicazione: (2025)
Enhanced Kalman with Adaptive Appearance Motion SORT for Grounded Generic Multiple Object Tracking
di: Anh, Duy Le Dinh, et al.
Pubblicazione: (2024)
di: Anh, Duy Le Dinh, et al.
Pubblicazione: (2024)
ActAlign: Zero-Shot Fine-Grained Video Classification via Language-Guided Sequence Alignment
di: Aghdam, Amir, et al.
Pubblicazione: (2025)
di: Aghdam, Amir, et al.
Pubblicazione: (2025)
Beyond Visual Understanding: Introducing PARROT-360V for Vision Language Model Benchmarking
di: Khurdula, Harsha Vardhan, et al.
Pubblicazione: (2024)
di: Khurdula, Harsha Vardhan, et al.
Pubblicazione: (2024)
Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images
di: Chen, Yuangong, et al.
Pubblicazione: (2026)
di: Chen, Yuangong, et al.
Pubblicazione: (2026)
Emo3D: Metric and Benchmarking Dataset for 3D Facial Expression Generation from Emotion Description
di: Dehghani, Mahshid, et al.
Pubblicazione: (2024)
di: Dehghani, Mahshid, et al.
Pubblicazione: (2024)
GroundCap: A Visually Grounded Image Captioning Dataset
di: Oliveira, Daniel A. P., et al.
Pubblicazione: (2025)
di: Oliveira, Daniel A. P., et al.
Pubblicazione: (2025)
Relative Drawing Identification Complexity is Invariant to Modality in Vision-Language Models
di: Freitas, Diogo, et al.
Pubblicazione: (2025)
di: Freitas, Diogo, et al.
Pubblicazione: (2025)
More Than Meets the Eye: Measuring the Semiotic Gap in Vision-Language Models via Semantic Anchorage
di: He, Wei
Pubblicazione: (2026)
di: He, Wei
Pubblicazione: (2026)
Story Generation from Visual Inputs: Techniques, Related Tasks, and Challenges
di: Oliveira, Daniel A. P., et al.
Pubblicazione: (2024)
di: Oliveira, Daniel A. P., et al.
Pubblicazione: (2024)
StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation
di: Oliveira, Daniel A. P., et al.
Pubblicazione: (2025)
di: Oliveira, Daniel A. P., et al.
Pubblicazione: (2025)
Towards Explainable Fake Image Detection with Multi-Modal Large Language Models
di: Ji, Yikun, et al.
Pubblicazione: (2025)
di: Ji, Yikun, et al.
Pubblicazione: (2025)
HuMoCon: Concept Discovery for Human Motion Understanding
di: Fang, Qihang, et al.
Pubblicazione: (2025)
di: Fang, Qihang, et al.
Pubblicazione: (2025)
Step-CoT: Stepwise Visual Chain-of-Thought for Medical Visual Question Answering
di: Fan, Lin, et al.
Pubblicazione: (2026)
di: Fan, Lin, et al.
Pubblicazione: (2026)
VGA: Vision GUI Assistant -- Minimizing Hallucinations through Image-Centric Fine-Tuning
di: Meng, Ziyang, et al.
Pubblicazione: (2024)
di: Meng, Ziyang, et al.
Pubblicazione: (2024)
Enhancing Sports Strategy with Video Analytics and Data Mining: Assessing the effectiveness of Multimodal LLMs in tennis video analysis
di: Teo, Charlton
Pubblicazione: (2025)
di: Teo, Charlton
Pubblicazione: (2025)
InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information
di: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Pubblicazione: (2025)
di: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Pubblicazione: (2025)
VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena
di: Parcalabescu, Letitia, et al.
Pubblicazione: (2021)
di: Parcalabescu, Letitia, et al.
Pubblicazione: (2021)
MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue
di: Deichler, Anna, et al.
Pubblicazione: (2026)
di: Deichler, Anna, et al.
Pubblicazione: (2026)
WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents
di: Liu, Bingnan, et al.
Pubblicazione: (2026)
di: Liu, Bingnan, et al.
Pubblicazione: (2026)
Analyzing Quality, Bias, and Performance in Text-to-Image Generative Models
di: Masrourisaadat, Nila, et al.
Pubblicazione: (2024)
di: Masrourisaadat, Nila, et al.
Pubblicazione: (2024)
ZeShot-VQA: Zero-Shot Visual Question Answering Framework with Answer Mapping for Natural Disaster Damage Assessment
di: Karimi, Ehsan, et al.
Pubblicazione: (2025)
di: Karimi, Ehsan, et al.
Pubblicazione: (2025)
A Surveillance Based Interactive Robot
di: Kavimandan, Kshitij, et al.
Pubblicazione: (2025)
di: Kavimandan, Kshitij, et al.
Pubblicazione: (2025)
Tri-VQA: Triangular Reasoning Medical Visual Question Answering for Multi-Attribute Analysis
di: Fan, Lin, et al.
Pubblicazione: (2024)
di: Fan, Lin, et al.
Pubblicazione: (2024)
VLEU: a Method for Automatic Evaluation for Generalizability of Text-to-Image Models
di: Cao, Jingtao, et al.
Pubblicazione: (2024)
di: Cao, Jingtao, et al.
Pubblicazione: (2024)
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
di: Dai, Song, et al.
Pubblicazione: (2025)
di: Dai, Song, et al.
Pubblicazione: (2025)
On the development of an AI performance and behavioural measures for teaching and classroom management
di: Niculescu, Andreea I., et al.
Pubblicazione: (2025)
di: Niculescu, Andreea I., et al.
Pubblicazione: (2025)
A Multi-Modal Deep Learning Based Approach for House Price Prediction
di: Hasan, Md Hasebul, et al.
Pubblicazione: (2024)
di: Hasan, Md Hasebul, et al.
Pubblicazione: (2024)
Semantic Leakage from Image Embeddings
di: Chen, Yiyi, et al.
Pubblicazione: (2026)
di: Chen, Yiyi, et al.
Pubblicazione: (2026)
CORDIAL: Can Multimodal Large Language Models Effectively Understand Coherence Relationships?
di: Ramakrishnan, Aashish Anantha, et al.
Pubblicazione: (2025)
di: Ramakrishnan, Aashish Anantha, et al.
Pubblicazione: (2025)
MM-SHAP: A Performance-agnostic Metric for Measuring Multimodal Contributions in Vision and Language Models & Tasks
di: Parcalabescu, Letitia, et al.
Pubblicazione: (2022)
di: Parcalabescu, Letitia, et al.
Pubblicazione: (2022)
TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning
di: Sanders, Kate, et al.
Pubblicazione: (2024)
di: Sanders, Kate, et al.
Pubblicazione: (2024)
On the Limitations of Vision-Language Models in Understanding Image Transforms
di: Anis, Ahmad Mustafa, et al.
Pubblicazione: (2025)
di: Anis, Ahmad Mustafa, et al.
Pubblicazione: (2025)
DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction
di: Hou, Zhiyi, et al.
Pubblicazione: (2025)
di: Hou, Zhiyi, et al.
Pubblicazione: (2025)
SpatialMath: Spatial Comprehension-Infused Symbolic Reasoning for Mathematical Problem-Solving
di: Bajpai, Ashutosh, et al.
Pubblicazione: (2026)
di: Bajpai, Ashutosh, et al.
Pubblicazione: (2026)
MDA: An Interpretable and Scalable Multi-Modal Fusion under Missing Modalities and Intrinsic Noise Conditions
di: Fan, Lin, et al.
Pubblicazione: (2024)
di: Fan, Lin, et al.
Pubblicazione: (2024)
Aligning by Misaligning: Boundary-aware Curriculum Learning for Multimodal Alignment
di: Ye, Hua, et al.
Pubblicazione: (2025)
di: Ye, Hua, et al.
Pubblicazione: (2025)
Text-to-Events: Synthetic Event Camera Streams from Conditional Text Input
di: Ott, Joachim, et al.
Pubblicazione: (2024)
di: Ott, Joachim, et al.
Pubblicazione: (2024)
Documenti analoghi
-
VidNum-1.4K: A Comprehensive Benchmark for Video-based Numerical Reasoning
di: Cui, Shaoyang, et al.
Pubblicazione: (2026) -
StoryMovie: A Dataset for Semantic Alignment of Visual Stories with Movie Scripts and Subtitles
di: Oliveira, Daniel, et al.
Pubblicazione: (2026) -
ReSpace: Text-Driven Autoregressive 3D Indoor Scene Synthesis and Editing
di: Bucher, Martin JJ., et al.
Pubblicazione: (2025) -
Enhanced Kalman with Adaptive Appearance Motion SORT for Grounded Generic Multiple Object Tracking
di: Anh, Duy Le Dinh, et al.
Pubblicazione: (2024) -
ActAlign: Zero-Shot Fine-Grained Video Classification via Language-Guided Sequence Alignment
di: Aghdam, Amir, et al.
Pubblicazione: (2025)