SAW-Bench: Learning Situated Awareness in the Real World

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Chuhan, Han, Rilyn, Hsu, Joy, Liang, Yongyuan, Dhawan, Rajiv, Wu, Jiajun, Yang, Ming-Hsuan, Wang, Xin Eric
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918530836856832
author Li, Chuhan
Han, Rilyn
Hsu, Joy
Liang, Yongyuan
Dhawan, Rajiv
Wu, Jiajun
Yang, Ming-Hsuan
Wang, Xin Eric
author_facet Li, Chuhan
Han, Rilyn
Hsu, Joy
Liang, Yongyuan
Dhawan, Rajiv
Wu, Jiajun
Yang, Ming-Hsuan
Wang, Xin Eric
contents A core aspect of human perception is situated awareness, the ability to relate ourselves to the surrounding physical environment and reason over possible actions in context. However, most existing benchmarks for multimodal foundation models (MFMs) emphasize environment-centric spatial relations (relations among objects in a scene), while largely overlooking observer-centric relationships that require reasoning relative to agent's viewpoint, pose, and motion. To bridge this gap, we introduce SAW-Bench (Situated Awareness in the Real World), a novel benchmark for evaluating egocentric situated awareness using real-world videos. SAW-Bench comprises 786 self-recorded videos captured with Ray-Ban Meta (Gen 2) smart glasses spanning diverse indoor and outdoor environments, and over 2,071 human-annotated question-answer pairs. It probes a model's observer-centric understanding with six different awareness tasks. Our comprehensive evaluation reveals a human-model performance gap of 37.66%, even with the best-performing MFM, Gemini 3 Flash. Beyond this gap, our in-depth analysis uncovers several notable findings; for example, while models can exploit partial geometric cues in egocentric videos, they often fail to infer a coherent camera geometry, leading to systematic spatial reasoning errors. We position SAW-Bench as a benchmark for situated spatial intelligence, moving beyond passive observation to understanding physically grounded, observer-centric dynamics.
format Preprint
id arxiv_https___arxiv_org_abs_2602_16682
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SAW-Bench: Learning Situated Awareness in the Real World
Li, Chuhan
Han, Rilyn
Hsu, Joy
Liang, Yongyuan
Dhawan, Rajiv
Wu, Jiajun
Yang, Ming-Hsuan
Wang, Xin Eric
Computer Vision and Pattern Recognition
A core aspect of human perception is situated awareness, the ability to relate ourselves to the surrounding physical environment and reason over possible actions in context. However, most existing benchmarks for multimodal foundation models (MFMs) emphasize environment-centric spatial relations (relations among objects in a scene), while largely overlooking observer-centric relationships that require reasoning relative to agent's viewpoint, pose, and motion. To bridge this gap, we introduce SAW-Bench (Situated Awareness in the Real World), a novel benchmark for evaluating egocentric situated awareness using real-world videos. SAW-Bench comprises 786 self-recorded videos captured with Ray-Ban Meta (Gen 2) smart glasses spanning diverse indoor and outdoor environments, and over 2,071 human-annotated question-answer pairs. It probes a model's observer-centric understanding with six different awareness tasks. Our comprehensive evaluation reveals a human-model performance gap of 37.66%, even with the best-performing MFM, Gemini 3 Flash. Beyond this gap, our in-depth analysis uncovers several notable findings; for example, while models can exploit partial geometric cues in egocentric videos, they often fail to infer a coherent camera geometry, leading to systematic spatial reasoning errors. We position SAW-Bench as a benchmark for situated spatial intelligence, moving beyond passive observation to understanding physically grounded, observer-centric dynamics.
title SAW-Bench: Learning Situated Awareness in the Real World
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.16682