Spot The Ball: A Benchmark for Visual Social Inference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Balamurugan, Neha, Wu, Sarah, Chun, Adam, Gaw, Gabe, Eyzaguirre, Cristobal, Gerstenberg, Tobias
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915626478469120
author Balamurugan, Neha
Wu, Sarah
Chun, Adam
Gaw, Gabe
Eyzaguirre, Cristobal
Gerstenberg, Tobias
author_facet Balamurugan, Neha
Wu, Sarah
Chun, Adam
Gaw, Gabe
Eyzaguirre, Cristobal
Gerstenberg, Tobias
contents Humans excel at visual social inference, the ability to infer hidden elements of a scene from subtle behavioral cues such as other people's gaze, pose, and orientation. This ability drives everyday social reasoning in humans and is critical for developing more human-like AI agents. We introduce Spot The Ball, a challenging benchmark for evaluating visual social inference in vision-language models (VLMs) using sports as a test domain. The task is to localize a removed sports ball from soccer, basketball, and volleyball images. We present a curated evaluation set with human baselines and a scalable pipeline for generating additional test items. We evaluate four state-of-the-art VLMs (Gemini, GPT, LLaMA, Qwen) using three prompting strategies, finding that humans are consistently two to three times more accurate (20-34%) than models ($\leq$ 17%) across all sports. Our analyses show that models rely on superficial spatial heuristics--such as guessing near the image center or nearby players--while humans leverage social cues like gaze direction and body pose. These findings reveal a persistent human-model gap in visual social reasoning and underscore the need for architectures that explicitly encode structured behavioral cues to achieve robust, human-like inference.
format Preprint
id arxiv_https___arxiv_org_abs_2511_00261
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Spot The Ball: A Benchmark for Visual Social Inference
Balamurugan, Neha
Wu, Sarah
Chun, Adam
Gaw, Gabe
Eyzaguirre, Cristobal
Gerstenberg, Tobias
Computer Vision and Pattern Recognition
Human-Computer Interaction
Humans excel at visual social inference, the ability to infer hidden elements of a scene from subtle behavioral cues such as other people's gaze, pose, and orientation. This ability drives everyday social reasoning in humans and is critical for developing more human-like AI agents. We introduce Spot The Ball, a challenging benchmark for evaluating visual social inference in vision-language models (VLMs) using sports as a test domain. The task is to localize a removed sports ball from soccer, basketball, and volleyball images. We present a curated evaluation set with human baselines and a scalable pipeline for generating additional test items. We evaluate four state-of-the-art VLMs (Gemini, GPT, LLaMA, Qwen) using three prompting strategies, finding that humans are consistently two to three times more accurate (20-34%) than models ($\leq$ 17%) across all sports. Our analyses show that models rely on superficial spatial heuristics--such as guessing near the image center or nearby players--while humans leverage social cues like gaze direction and body pose. These findings reveal a persistent human-model gap in visual social reasoning and underscore the need for architectures that explicitly encode structured behavioral cues to achieve robust, human-like inference.
title Spot The Ball: A Benchmark for Visual Social Inference
topic Computer Vision and Pattern Recognition
Human-Computer Interaction
url https://arxiv.org/abs/2511.00261