PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Qiran, Wang, Yuheng, Yang, Runde, Wu, Lin, Fan, Jingru, Yao, Shu, Zhang, Jie, Zhou, Tianle, Li, Huatao, Shi, Ruijie, Li, Yihan, Qian, Chen
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914579479527424
author Zhang, Qiran
Wang, Yuheng
Yang, Runde
Wu, Lin
Fan, Jingru
Yao, Shu
Zhang, Jie
Zhou, Tianle
Li, Huatao
Shi, Ruijie
Li, Yihan
Qian, Chen
author_facet Zhang, Qiran
Wang, Yuheng
Yang, Runde
Wu, Lin
Fan, Jingru
Yao, Shu
Zhang, Jie
Zhou, Tianle
Li, Huatao
Shi, Ruijie
Li, Yihan
Qian, Chen
contents Programmatic video generation through code offers geometric precision and temporal coherence beyond pixel-level diffusion models, yet rigorously evaluating whether language models can produce spatially correct animated outputs remains an open problem. We introduce PRISM, a large-scale benchmark of 10,372 human-calibrated instruction-code pairs (20 times larger than prior programmatic video generation benchmarks), grounded in real-world knowledge visualization scenarios across English and Chinese and spanning 437 subject categories. We further propose a funnel-style evaluation framework with four complementary metrics: Code-Level Reliability for executability, Spatial Reasoning for layout correctness over full animation sequences, and Prompt-Aware Dynamic Visual Complexity (PADVC) and Temporal Density (TD) for diagnosing dynamic expression and temporal activity. Systematic evaluation of seven mainstream LLMs reveals a striking Execution-Spatial Gap: the average drop from execution success rate to spatial pass rate is approximately 41%, showing that runnable code does not necessarily yield spatially coherent visual output. These findings show that programmatic video generation evaluation should go beyond executability. PRISM provides a principled benchmark for advancing spatially coherent code generation.
format Preprint
id arxiv_https___arxiv_org_abs_2605_19382
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning
Zhang, Qiran
Wang, Yuheng
Yang, Runde
Wu, Lin
Fan, Jingru
Yao, Shu
Zhang, Jie
Zhou, Tianle
Li, Huatao
Shi, Ruijie
Li, Yihan
Qian, Chen
Artificial Intelligence
Programmatic video generation through code offers geometric precision and temporal coherence beyond pixel-level diffusion models, yet rigorously evaluating whether language models can produce spatially correct animated outputs remains an open problem. We introduce PRISM, a large-scale benchmark of 10,372 human-calibrated instruction-code pairs (20 times larger than prior programmatic video generation benchmarks), grounded in real-world knowledge visualization scenarios across English and Chinese and spanning 437 subject categories. We further propose a funnel-style evaluation framework with four complementary metrics: Code-Level Reliability for executability, Spatial Reasoning for layout correctness over full animation sequences, and Prompt-Aware Dynamic Visual Complexity (PADVC) and Temporal Density (TD) for diagnosing dynamic expression and temporal activity. Systematic evaluation of seven mainstream LLMs reveals a striking Execution-Spatial Gap: the average drop from execution success rate to spatial pass rate is approximately 41%, showing that runnable code does not necessarily yield spatially coherent visual output. These findings show that programmatic video generation evaluation should go beyond executability. PRISM provides a principled benchmark for advancing spatially coherent code generation.
title PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning
topic Artificial Intelligence
url https://arxiv.org/abs/2605.19382