m2sv: A Scalable Benchmark for Map-to-Street-View Spatial Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shin, Yosub, Buriek, Michael, Molybog, Igor |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond Audio and Pose: A General-Purpose Framework for Video Synchronization
von: Shin, Yosub, et al.
Veröffentlicht: (2025)
von: Shin, Yosub, et al.
Veröffentlicht: (2025)
What Matters in Data Curation for Multimodal Reasoning? Insights from the DCVLR Challenge
von: Shin, Yosub, et al.
Veröffentlicht: (2026)
von: Shin, Yosub, et al.
Veröffentlicht: (2026)
Combining Deep Learning and Street View Imagery to Map Smallholder Crop Types
von: Soler, Jordi Laguarta, et al.
Veröffentlicht: (2023)
von: Soler, Jordi Laguarta, et al.
Veröffentlicht: (2023)
Cross-View Geolocalization and Disaster Mapping with Street-View and VHR Satellite Imagery: A Case Study of Hurricane IAN
von: Li, Hao, et al.
Veröffentlicht: (2024)
von: Li, Hao, et al.
Veröffentlicht: (2024)
Unsupervised Urban Land Use Mapping with Street View Contrastive Clustering and a Geographical Prior
von: Che, Lin, et al.
Veröffentlicht: (2025)
von: Che, Lin, et al.
Veröffentlicht: (2025)
Learning Street View Representations with Spatiotemporal Contrast
von: Li, Yong, et al.
Veröffentlicht: (2025)
von: Li, Yong, et al.
Veröffentlicht: (2025)
CoV: Chain-of-View Prompting for Spatial Reasoning
von: Zhao, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhao, Haoyu, et al.
Veröffentlicht: (2026)
Bird Eye-View to Street-View: A Survey
von: Bajbaa, Khawlah, et al.
Veröffentlicht: (2024)
von: Bajbaa, Khawlah, et al.
Veröffentlicht: (2024)
Map2Thought: Explicit 3D Spatial Reasoning via Metric Cognitive Maps
von: Gao, Xiangjun, et al.
Veröffentlicht: (2026)
von: Gao, Xiangjun, et al.
Veröffentlicht: (2026)
MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning
von: Jiang, Yulun, et al.
Veröffentlicht: (2025)
von: Jiang, Yulun, et al.
Veröffentlicht: (2025)
Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Multimodal Models
von: Wang, Xingrui, et al.
Veröffentlicht: (2025)
von: Wang, Xingrui, et al.
Veröffentlicht: (2025)
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence
von: Liu, Chonghan, et al.
Veröffentlicht: (2025)
von: Liu, Chonghan, et al.
Veröffentlicht: (2025)
Examining the Commitments and Difficulties Inherent in Multimodal Foundation Models for Street View Imagery
von: Yang, Zhenyuan, et al.
Veröffentlicht: (2024)
von: Yang, Zhenyuan, et al.
Veröffentlicht: (2024)
OpenStreetView-5M: The Many Roads to Global Visual Geolocation
von: Astruc, Guillaume, et al.
Veröffentlicht: (2024)
von: Astruc, Guillaume, et al.
Veröffentlicht: (2024)
MagicDrive: Street View Generation with Diverse 3D Geometry Control
von: Gao, Ruiyuan, et al.
Veröffentlicht: (2023)
von: Gao, Ruiyuan, et al.
Veröffentlicht: (2023)
EarthSpatialBench: Benchmarking Spatial Reasoning Capabilities of Multimodal LLMs on Earth Imagery
von: Xu, Zelin, et al.
Veröffentlicht: (2026)
von: Xu, Zelin, et al.
Veröffentlicht: (2026)
MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model
von: Lee, Youngwan, et al.
Veröffentlicht: (2026)
von: Lee, Youngwan, et al.
Veröffentlicht: (2026)
Mapping License Plate Recoverability Under Extreme Viewing Angles for Oppor-tunistic Urban Sensing
von: Adamenko, Igor, et al.
Veröffentlicht: (2026)
von: Adamenko, Igor, et al.
Veröffentlicht: (2026)
NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
von: Tian, Kexin, et al.
Veröffentlicht: (2025)
von: Tian, Kexin, et al.
Veröffentlicht: (2025)
Oitijjo-3D: Generative AI Framework for Rapid 3D Heritage Reconstruction from Street View Imagery
von: Ope, Momen Khandoker, et al.
Veröffentlicht: (2025)
von: Ope, Momen Khandoker, et al.
Veröffentlicht: (2025)
CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark
von: Wang, Wei, et al.
Veröffentlicht: (2026)
von: Wang, Wei, et al.
Veröffentlicht: (2026)
Leveraging Multimodal LLMs for Built Environment and Housing Attribute Assessment from Street-View Imagery
von: Yao, Siyuan, et al.
Veröffentlicht: (2026)
von: Yao, Siyuan, et al.
Veröffentlicht: (2026)
An Integrated Causal Inference Framework for Traffic Safety Modeling with Semantic Street-View Visual Features
von: Sun, Lishan, et al.
Veröffentlicht: (2026)
von: Sun, Lishan, et al.
Veröffentlicht: (2026)
Cube Bench: A Benchmark for Spatial Visual Reasoning in MLLMs
von: Anand, Dhruv, et al.
Veröffentlicht: (2025)
von: Anand, Dhruv, et al.
Veröffentlicht: (2025)
Tag Map: A Text-Based Map for Spatial Reasoning and Navigation with Large Language Models
von: Zhang, Mike, et al.
Veröffentlicht: (2024)
von: Zhang, Mike, et al.
Veröffentlicht: (2024)
Robust Vehicle Localization and Tracking in Rain using Street Maps
von: Tan, Yu Xiang, et al.
Veröffentlicht: (2024)
von: Tan, Yu Xiang, et al.
Veröffentlicht: (2024)
OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
von: Jia, Mengdi, et al.
Veröffentlicht: (2025)
von: Jia, Mengdi, et al.
Veröffentlicht: (2025)
Satellite-to-Street: Synthesizing Post-Disaster Views from Satellite Imagery via Generative Vision Models
von: Yang, Yifan, et al.
Veröffentlicht: (2026)
von: Yang, Yifan, et al.
Veröffentlicht: (2026)
From Bird's-Eye to Street View: Crafting Diverse and Condition-Aligned Images with Latent Diffusion Model
von: Xu, Xiaojie, et al.
Veröffentlicht: (2024)
von: Xu, Xiaojie, et al.
Veröffentlicht: (2024)
MagicDrive3D: Controllable 3D Generation for Any-View Rendering in Street Scenes
von: Gao, Ruiyuan, et al.
Veröffentlicht: (2024)
von: Gao, Ruiyuan, et al.
Veröffentlicht: (2024)
From Street to Orbit: Training-Free Cross-View Retrieval via Location Semantics and LLM Guidance
von: Min, Jeongho, et al.
Veröffentlicht: (2025)
von: Min, Jeongho, et al.
Veröffentlicht: (2025)
Paved or unpaved? A Deep Learning derived Road Surface Global Dataset from Mapillary Street-View Imagery
von: Randhawa, Sukanya, et al.
Veröffentlicht: (2024)
von: Randhawa, Sukanya, et al.
Veröffentlicht: (2024)
BuildingView: Constructing Urban Building Exteriors Databases with Street View Imagery and Multimodal Large Language Mode
von: Li, Zongrong, et al.
Veröffentlicht: (2024)
von: Li, Zongrong, et al.
Veröffentlicht: (2024)
VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street View
von: Schumann, Raphael, et al.
Veröffentlicht: (2023)
von: Schumann, Raphael, et al.
Veröffentlicht: (2023)
MapTrace: Scalable Data Generation for Route Tracing on Maps
von: Panagopoulou, Artemis, et al.
Veröffentlicht: (2025)
von: Panagopoulou, Artemis, et al.
Veröffentlicht: (2025)
iVISPAR -- An Interactive Visual-Spatial Reasoning Benchmark for VLMs
von: Mayer, Julius, et al.
Veröffentlicht: (2025)
von: Mayer, Julius, et al.
Veröffentlicht: (2025)
WaymoQA: A Multi-View Visual Question Answering Dataset for Safety-Critical Reasoning in Autonomous Driving
von: Yu, Seungjun, et al.
Veröffentlicht: (2025)
von: Yu, Seungjun, et al.
Veröffentlicht: (2025)
DreamDrive: Generative 4D Scene Modeling from Street View Images
von: Mao, Jiageng, et al.
Veröffentlicht: (2024)
von: Mao, Jiageng, et al.
Veröffentlicht: (2024)
CityCube: Benchmarking Cross-view Spatial Reasoning on Vision-Language Models in Urban Environments
von: Xu, Haotian, et al.
Veröffentlicht: (2026)
von: Xu, Haotian, et al.
Veröffentlicht: (2026)
Memory-Scalable and Simplified Functional Map Learning
von: Magnet, Robin, et al.
Veröffentlicht: (2024)
von: Magnet, Robin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Beyond Audio and Pose: A General-Purpose Framework for Video Synchronization
von: Shin, Yosub, et al.
Veröffentlicht: (2025) -
What Matters in Data Curation for Multimodal Reasoning? Insights from the DCVLR Challenge
von: Shin, Yosub, et al.
Veröffentlicht: (2026) -
Combining Deep Learning and Street View Imagery to Map Smallholder Crop Types
von: Soler, Jordi Laguarta, et al.
Veröffentlicht: (2023) -
Cross-View Geolocalization and Disaster Mapping with Street-View and VHR Satellite Imagery: A Case Study of Hurricane IAN
von: Li, Hao, et al.
Veröffentlicht: (2024) -
Unsupervised Urban Land Use Mapping with Street View Contrastive Clustering and a Geographical Prior
von: Che, Lin, et al.
Veröffentlicht: (2025)