PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914579479527424 |
|---|---|
| author | Zhang, Qiran Wang, Yuheng Yang, Runde Wu, Lin Fan, Jingru Yao, Shu Zhang, Jie Zhou, Tianle Li, Huatao Shi, Ruijie Li, Yihan Qian, Chen |
| author_facet | Zhang, Qiran Wang, Yuheng Yang, Runde Wu, Lin Fan, Jingru Yao, Shu Zhang, Jie Zhou, Tianle Li, Huatao Shi, Ruijie Li, Yihan Qian, Chen |
| contents | Programmatic video generation through code offers geometric precision and temporal coherence beyond pixel-level diffusion models, yet rigorously evaluating whether language models can produce spatially correct animated outputs remains an open problem. We introduce PRISM, a large-scale benchmark of 10,372 human-calibrated instruction-code pairs (20 times larger than prior programmatic video generation benchmarks), grounded in real-world knowledge visualization scenarios across English and Chinese and spanning 437 subject categories. We further propose a funnel-style evaluation framework with four complementary metrics: Code-Level Reliability for executability, Spatial Reasoning for layout correctness over full animation sequences, and Prompt-Aware Dynamic Visual Complexity (PADVC) and Temporal Density (TD) for diagnosing dynamic expression and temporal activity. Systematic evaluation of seven mainstream LLMs reveals a striking Execution-Spatial Gap: the average drop from execution success rate to spatial pass rate is approximately 41%, showing that runnable code does not necessarily yield spatially coherent visual output. These findings show that programmatic video generation evaluation should go beyond executability. PRISM provides a principled benchmark for advancing spatially coherent code generation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_19382 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning Zhang, Qiran Wang, Yuheng Yang, Runde Wu, Lin Fan, Jingru Yao, Shu Zhang, Jie Zhou, Tianle Li, Huatao Shi, Ruijie Li, Yihan Qian, Chen Artificial Intelligence Programmatic video generation through code offers geometric precision and temporal coherence beyond pixel-level diffusion models, yet rigorously evaluating whether language models can produce spatially correct animated outputs remains an open problem. We introduce PRISM, a large-scale benchmark of 10,372 human-calibrated instruction-code pairs (20 times larger than prior programmatic video generation benchmarks), grounded in real-world knowledge visualization scenarios across English and Chinese and spanning 437 subject categories. We further propose a funnel-style evaluation framework with four complementary metrics: Code-Level Reliability for executability, Spatial Reasoning for layout correctness over full animation sequences, and Prompt-Aware Dynamic Visual Complexity (PADVC) and Temporal Density (TD) for diagnosing dynamic expression and temporal activity. Systematic evaluation of seven mainstream LLMs reveals a striking Execution-Spatial Gap: the average drop from execution success rate to spatial pass rate is approximately 41%, showing that runnable code does not necessarily yield spatially coherent visual output. These findings show that programmatic video generation evaluation should go beyond executability. PRISM provides a principled benchmark for advancing spatially coherent code generation. |
| title | PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2605.19382 |