TimeSpot: Benchmarking Geo-Temporal Understanding in Vision-Language Models in Real-World Settings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wasi, Azmine Toushik, Ridoy, Shahriyar Zaman, Tonmoy, Koushik Ahamed, Tshering, Kinga, Hasan, S. M. Muhtasimul, Faisal, Wahid, Mohiuddin, Tasnim, Parvez, Md Rizwan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918520638406656
author Wasi, Azmine Toushik
Ridoy, Shahriyar Zaman
Tonmoy, Koushik Ahamed
Tshering, Kinga
Hasan, S. M. Muhtasimul
Faisal, Wahid
Mohiuddin, Tasnim
Parvez, Md Rizwan
author_facet Wasi, Azmine Toushik
Ridoy, Shahriyar Zaman
Tonmoy, Koushik Ahamed
Tshering, Kinga
Hasan, S. M. Muhtasimul
Faisal, Wahid
Mohiuddin, Tasnim
Parvez, Md Rizwan
contents Geo-temporal understanding, the ability to infer location, time, and contextual properties from visual input alone, underpins applications such as disaster management, traffic planning, embodied navigation, world modeling, and geography education. Although recent vision-language models (VLMs) have advanced image geo-localization using cues like landmarks and road signs, their ability to reason about temporal signals and physically grounded spatial cues remains limited. To address this gap, we introduce TimeSpot, a benchmark for evaluating real-world geo-temporal reasoning in VLMs. TimeSpot comprises 1,455 ground-level images from 80 countries and requires structured prediction of temporal attributes (season, month, time of day, daylight phase) and geographic attributes (continent, country, climate zone, environment type, latitude-longitude) directly from visual evidence. It also includes spatial-temporal reasoning tasks that test physical plausibility under real-world uncertainty. Evaluations of state-of-the-art open- and closed-source VLMs show low performance, particularly for temporal inference. While supervised fine-tuning yields improvements, results remain insufficient, highlighting the need for new methods to achieve robust, physically grounded geo-temporal understanding TimeSpot is available at: https://TimeSpot-GT.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2603_06687
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TimeSpot: Benchmarking Geo-Temporal Understanding in Vision-Language Models in Real-World Settings
Wasi, Azmine Toushik
Ridoy, Shahriyar Zaman
Tonmoy, Koushik Ahamed
Tshering, Kinga
Hasan, S. M. Muhtasimul
Faisal, Wahid
Mohiuddin, Tasnim
Parvez, Md Rizwan
Computer Vision and Pattern Recognition
Computation and Language
Emerging Technologies
Multimedia
Robotics
Geo-temporal understanding, the ability to infer location, time, and contextual properties from visual input alone, underpins applications such as disaster management, traffic planning, embodied navigation, world modeling, and geography education. Although recent vision-language models (VLMs) have advanced image geo-localization using cues like landmarks and road signs, their ability to reason about temporal signals and physically grounded spatial cues remains limited. To address this gap, we introduce TimeSpot, a benchmark for evaluating real-world geo-temporal reasoning in VLMs. TimeSpot comprises 1,455 ground-level images from 80 countries and requires structured prediction of temporal attributes (season, month, time of day, daylight phase) and geographic attributes (continent, country, climate zone, environment type, latitude-longitude) directly from visual evidence. It also includes spatial-temporal reasoning tasks that test physical plausibility under real-world uncertainty. Evaluations of state-of-the-art open- and closed-source VLMs show low performance, particularly for temporal inference. While supervised fine-tuning yields improvements, results remain insufficient, highlighting the need for new methods to achieve robust, physically grounded geo-temporal understanding TimeSpot is available at: https://TimeSpot-GT.github.io.
title TimeSpot: Benchmarking Geo-Temporal Understanding in Vision-Language Models in Real-World Settings
topic Computer Vision and Pattern Recognition
Computation and Language
Emerging Technologies
Multimedia
Robotics
url https://arxiv.org/abs/2603.06687