m2sv: A Scalable Benchmark for Map-to-Street-View Spatial Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Shin, Yosub, Buriek, Michael, Molybog, Igor |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Audio and Pose: A General-Purpose Framework for Video Synchronization
by: Shin, Yosub, et al.
Published: (2025)
by: Shin, Yosub, et al.
Published: (2025)
What Matters in Data Curation for Multimodal Reasoning? Insights from the DCVLR Challenge
by: Shin, Yosub, et al.
Published: (2026)
by: Shin, Yosub, et al.
Published: (2026)
Combining Deep Learning and Street View Imagery to Map Smallholder Crop Types
by: Soler, Jordi Laguarta, et al.
Published: (2023)
by: Soler, Jordi Laguarta, et al.
Published: (2023)
Cross-View Geolocalization and Disaster Mapping with Street-View and VHR Satellite Imagery: A Case Study of Hurricane IAN
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Unsupervised Urban Land Use Mapping with Street View Contrastive Clustering and a Geographical Prior
by: Che, Lin, et al.
Published: (2025)
by: Che, Lin, et al.
Published: (2025)
Learning Street View Representations with Spatiotemporal Contrast
by: Li, Yong, et al.
Published: (2025)
by: Li, Yong, et al.
Published: (2025)
CoV: Chain-of-View Prompting for Spatial Reasoning
by: Zhao, Haoyu, et al.
Published: (2026)
by: Zhao, Haoyu, et al.
Published: (2026)
Bird Eye-View to Street-View: A Survey
by: Bajbaa, Khawlah, et al.
Published: (2024)
by: Bajbaa, Khawlah, et al.
Published: (2024)
Map2Thought: Explicit 3D Spatial Reasoning via Metric Cognitive Maps
by: Gao, Xiangjun, et al.
Published: (2026)
by: Gao, Xiangjun, et al.
Published: (2026)
MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning
by: Jiang, Yulun, et al.
Published: (2025)
by: Jiang, Yulun, et al.
Published: (2025)
Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Multimodal Models
by: Wang, Xingrui, et al.
Published: (2025)
by: Wang, Xingrui, et al.
Published: (2025)
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence
by: Liu, Chonghan, et al.
Published: (2025)
by: Liu, Chonghan, et al.
Published: (2025)
Examining the Commitments and Difficulties Inherent in Multimodal Foundation Models for Street View Imagery
by: Yang, Zhenyuan, et al.
Published: (2024)
by: Yang, Zhenyuan, et al.
Published: (2024)
OpenStreetView-5M: The Many Roads to Global Visual Geolocation
by: Astruc, Guillaume, et al.
Published: (2024)
by: Astruc, Guillaume, et al.
Published: (2024)
MagicDrive: Street View Generation with Diverse 3D Geometry Control
by: Gao, Ruiyuan, et al.
Published: (2023)
by: Gao, Ruiyuan, et al.
Published: (2023)
EarthSpatialBench: Benchmarking Spatial Reasoning Capabilities of Multimodal LLMs on Earth Imagery
by: Xu, Zelin, et al.
Published: (2026)
by: Xu, Zelin, et al.
Published: (2026)
MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model
by: Lee, Youngwan, et al.
Published: (2026)
by: Lee, Youngwan, et al.
Published: (2026)
Mapping License Plate Recoverability Under Extreme Viewing Angles for Oppor-tunistic Urban Sensing
by: Adamenko, Igor, et al.
Published: (2026)
by: Adamenko, Igor, et al.
Published: (2026)
NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
by: Tian, Kexin, et al.
Published: (2025)
by: Tian, Kexin, et al.
Published: (2025)
Oitijjo-3D: Generative AI Framework for Rapid 3D Heritage Reconstruction from Street View Imagery
by: Ope, Momen Khandoker, et al.
Published: (2025)
by: Ope, Momen Khandoker, et al.
Published: (2025)
CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark
by: Wang, Wei, et al.
Published: (2026)
by: Wang, Wei, et al.
Published: (2026)
Leveraging Multimodal LLMs for Built Environment and Housing Attribute Assessment from Street-View Imagery
by: Yao, Siyuan, et al.
Published: (2026)
by: Yao, Siyuan, et al.
Published: (2026)
An Integrated Causal Inference Framework for Traffic Safety Modeling with Semantic Street-View Visual Features
by: Sun, Lishan, et al.
Published: (2026)
by: Sun, Lishan, et al.
Published: (2026)
Cube Bench: A Benchmark for Spatial Visual Reasoning in MLLMs
by: Anand, Dhruv, et al.
Published: (2025)
by: Anand, Dhruv, et al.
Published: (2025)
Tag Map: A Text-Based Map for Spatial Reasoning and Navigation with Large Language Models
by: Zhang, Mike, et al.
Published: (2024)
by: Zhang, Mike, et al.
Published: (2024)
Robust Vehicle Localization and Tracking in Rain using Street Maps
by: Tan, Yu Xiang, et al.
Published: (2024)
by: Tan, Yu Xiang, et al.
Published: (2024)
OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
by: Jia, Mengdi, et al.
Published: (2025)
by: Jia, Mengdi, et al.
Published: (2025)
Satellite-to-Street: Synthesizing Post-Disaster Views from Satellite Imagery via Generative Vision Models
by: Yang, Yifan, et al.
Published: (2026)
by: Yang, Yifan, et al.
Published: (2026)
From Bird's-Eye to Street View: Crafting Diverse and Condition-Aligned Images with Latent Diffusion Model
by: Xu, Xiaojie, et al.
Published: (2024)
by: Xu, Xiaojie, et al.
Published: (2024)
MagicDrive3D: Controllable 3D Generation for Any-View Rendering in Street Scenes
by: Gao, Ruiyuan, et al.
Published: (2024)
by: Gao, Ruiyuan, et al.
Published: (2024)
From Street to Orbit: Training-Free Cross-View Retrieval via Location Semantics and LLM Guidance
by: Min, Jeongho, et al.
Published: (2025)
by: Min, Jeongho, et al.
Published: (2025)
Paved or unpaved? A Deep Learning derived Road Surface Global Dataset from Mapillary Street-View Imagery
by: Randhawa, Sukanya, et al.
Published: (2024)
by: Randhawa, Sukanya, et al.
Published: (2024)
BuildingView: Constructing Urban Building Exteriors Databases with Street View Imagery and Multimodal Large Language Mode
by: Li, Zongrong, et al.
Published: (2024)
by: Li, Zongrong, et al.
Published: (2024)
VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street View
by: Schumann, Raphael, et al.
Published: (2023)
by: Schumann, Raphael, et al.
Published: (2023)
MapTrace: Scalable Data Generation for Route Tracing on Maps
by: Panagopoulou, Artemis, et al.
Published: (2025)
by: Panagopoulou, Artemis, et al.
Published: (2025)
iVISPAR -- An Interactive Visual-Spatial Reasoning Benchmark for VLMs
by: Mayer, Julius, et al.
Published: (2025)
by: Mayer, Julius, et al.
Published: (2025)
WaymoQA: A Multi-View Visual Question Answering Dataset for Safety-Critical Reasoning in Autonomous Driving
by: Yu, Seungjun, et al.
Published: (2025)
by: Yu, Seungjun, et al.
Published: (2025)
DreamDrive: Generative 4D Scene Modeling from Street View Images
by: Mao, Jiageng, et al.
Published: (2024)
by: Mao, Jiageng, et al.
Published: (2024)
CityCube: Benchmarking Cross-view Spatial Reasoning on Vision-Language Models in Urban Environments
by: Xu, Haotian, et al.
Published: (2026)
by: Xu, Haotian, et al.
Published: (2026)
Memory-Scalable and Simplified Functional Map Learning
by: Magnet, Robin, et al.
Published: (2024)
by: Magnet, Robin, et al.
Published: (2024)
Similar Items
-
Beyond Audio and Pose: A General-Purpose Framework for Video Synchronization
by: Shin, Yosub, et al.
Published: (2025) -
What Matters in Data Curation for Multimodal Reasoning? Insights from the DCVLR Challenge
by: Shin, Yosub, et al.
Published: (2026) -
Combining Deep Learning and Street View Imagery to Map Smallholder Crop Types
by: Soler, Jordi Laguarta, et al.
Published: (2023) -
Cross-View Geolocalization and Disaster Mapping with Street-View and VHR Satellite Imagery: A Case Study of Hurricane IAN
by: Li, Hao, et al.
Published: (2024) -
Unsupervised Urban Land Use Mapping with Street View Contrastive Clustering and a Geographical Prior
by: Che, Lin, et al.
Published: (2025)