PRISM: A Multi-View Multi-Capability Retail Video Dataset for Embodied Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rouhi, Amirreza, Sakurikar, Parikshit, Reddy, Satya Sai, Menga, Narsimha, Govil, Anirudh, Chittajallu, Sri Harsha, Aggarwal, Rajat, Namboodiri, Anoop, Reddi, Sashi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912990614257664
author Rouhi, Amirreza
Sakurikar, Parikshit
Reddy, Satya Sai
Menga, Narsimha
Govil, Anirudh
Chittajallu, Sri Harsha
Aggarwal, Rajat
Namboodiri, Anoop
Reddi, Sashi
author_facet Rouhi, Amirreza
Sakurikar, Parikshit
Reddy, Satya Sai
Menga, Narsimha
Govil, Anirudh
Chittajallu, Sri Harsha
Aggarwal, Rajat
Namboodiri, Anoop
Reddi, Sashi
contents A critical gap exists between the general-purpose visual understanding of state-of-the-art physical AI models and the specialized perceptual demands of structured real-world deployment environments. We present PRISM, a 270K-sample multi-view video supervised fine-tuning (SFT) corpus for embodied vision-language-models (VLMs) in real-world retail environments. PRISM is motivated by a simple observation - physical AI systems fail not because of poor visual recognition, but because they do not understand space, physical dynamics and embodied action well enough to operate reliably in the world. To this end, PRISM is grounded in a novel three-dimensional knowledge ontology that spans spatial knowledge, temporal and physical knowledge, and embodied action knowledge. It covers 20+ capability probes across four evaluation dimensions - Embodied Reasoning (ER), Common Sense (CS), Spatial Perception (SP), and Intuitive Physics (IP), and to our knowledge, PRISM is the first dataset to instantiate all three knowledge dimensions within a single real-world deployment domain. The corpus captures data from egocentric, exocentric and 360° viewpoints across five supermarket locations and includes open-ended, chain-of-thought, and multiple-choice supervision. At 4 fps, PRISM spans approximately 11.8M video frames and approximately 730M tokens, placing it among the largest domain-specific video SFT corpora. Fine-tuning on PRISM reduces the error rate across all 20+ probes by 66.6% over the pre-trained baseline, with significant gains in embodied action understanding where the accuracy improves by 36.4%. Our results suggest that ontology-structured, domain specific SFT can meaningfully strengthen embodied VLMs for real-world settings. The PRISM dataset and more details are available at https://dreamvu.ai/prism
format Preprint
id arxiv_https___arxiv_org_abs_2603_29281
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PRISM: A Multi-View Multi-Capability Retail Video Dataset for Embodied Vision-Language Models
Rouhi, Amirreza
Sakurikar, Parikshit
Reddy, Satya Sai
Menga, Narsimha
Govil, Anirudh
Chittajallu, Sri Harsha
Aggarwal, Rajat
Namboodiri, Anoop
Reddi, Sashi
Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
A critical gap exists between the general-purpose visual understanding of state-of-the-art physical AI models and the specialized perceptual demands of structured real-world deployment environments. We present PRISM, a 270K-sample multi-view video supervised fine-tuning (SFT) corpus for embodied vision-language-models (VLMs) in real-world retail environments. PRISM is motivated by a simple observation - physical AI systems fail not because of poor visual recognition, but because they do not understand space, physical dynamics and embodied action well enough to operate reliably in the world. To this end, PRISM is grounded in a novel three-dimensional knowledge ontology that spans spatial knowledge, temporal and physical knowledge, and embodied action knowledge. It covers 20+ capability probes across four evaluation dimensions - Embodied Reasoning (ER), Common Sense (CS), Spatial Perception (SP), and Intuitive Physics (IP), and to our knowledge, PRISM is the first dataset to instantiate all three knowledge dimensions within a single real-world deployment domain. The corpus captures data from egocentric, exocentric and 360° viewpoints across five supermarket locations and includes open-ended, chain-of-thought, and multiple-choice supervision. At 4 fps, PRISM spans approximately 11.8M video frames and approximately 730M tokens, placing it among the largest domain-specific video SFT corpora. Fine-tuning on PRISM reduces the error rate across all 20+ probes by 66.6% over the pre-trained baseline, with significant gains in embodied action understanding where the accuracy improves by 36.4%. Our results suggest that ontology-structured, domain specific SFT can meaningfully strengthen embodied VLMs for real-world settings. The PRISM dataset and more details are available at https://dreamvu.ai/prism
title PRISM: A Multi-View Multi-Capability Retail Video Dataset for Embodied Vision-Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2603.29281