From Pixels to Places: A Systematic Benchmark for Evaluating Image Geolocalization Ability in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Lingyao, Yu, Runlong, Hu, Qikai, Li, Bowei, Deng, Min, Zhou, Yang, Jia, Xiaowei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Empowering LLM Agents with Geospatial Awareness: Toward Grounded Reasoning for Wildfire Response
by: Chen, Yiheng, et al.
Published: (2025)
by: Chen, Yiheng, et al.
Published: (2025)
LLMGeo: Benchmarking Large Language Models on Image Geolocation In-the-wild
by: Wang, Zhiqiang, et al.
Published: (2024)
by: Wang, Zhiqiang, et al.
Published: (2024)
When Earth Foundation Models Meet Diffusion: An Application to Land Surface Temperature Super-Resolution
by: Chen, Yiheng, et al.
Published: (2026)
by: Chen, Yiheng, et al.
Published: (2026)
Image-Based Geolocation Using Large Vision-Language Models
by: Liu, Yi, et al.
Published: (2024)
by: Liu, Yi, et al.
Published: (2024)
Learning to Wander: Improving the Global Image Geolocation Ability of LMMs via Actionable Reasoning
by: Zheng, Yushuo, et al.
Published: (2026)
by: Zheng, Yushuo, et al.
Published: (2026)
I2E: From Image Pixels to Actionable Interactive Environments for Text-Guided Image Editing
by: Yu, Jinghan, et al.
Published: (2026)
by: Yu, Jinghan, et al.
Published: (2026)
LLMs as World Models: Data-Driven and Human-Centered Pre-Event Simulation for Disaster Impact Assessment
by: Li, Lingyao, et al.
Published: (2025)
by: Li, Lingyao, et al.
Published: (2025)
VXP: Voxel-Cross-Pixel Large-scale Image-LiDAR Place Recognition
by: Li, Yun-Jin, et al.
Published: (2024)
by: Li, Yun-Jin, et al.
Published: (2024)
Enhance Image-to-Image Generation with LLaVA-generated Prompts
by: Ding, Zhicheng, et al.
Published: (2024)
by: Ding, Zhicheng, et al.
Published: (2024)
GeoRanker: Distance-Aware Ranking for Worldwide Image Geolocalization
by: Jia, Pengyue, et al.
Published: (2025)
by: Jia, Pengyue, et al.
Published: (2025)
Skill-Conditioned Visual Geolocation for Vision-Language Models
by: Yang, Chenjie, et al.
Published: (2026)
by: Yang, Chenjie, et al.
Published: (2026)
RAG for Geoscience: What We Expect, Gaps and Opportunities
by: Yu, Runlong, et al.
Published: (2025)
by: Yu, Runlong, et al.
Published: (2025)
LLM-as-a-Reviewer: Benchmarking Their Ability, Divergence, and Prompt Injection Resistance as Paper Reviewers
by: Li, Lingyao, et al.
Published: (2026)
by: Li, Lingyao, et al.
Published: (2026)
Can Large Multimodal Models Actively Recognize Faulty Inputs? A Systematic Evaluation Framework of Their Input Scrutiny Ability
by: Yang, Haiqi, et al.
Published: (2025)
by: Yang, Haiqi, et al.
Published: (2025)
From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models
by: Lin, Yihan, et al.
Published: (2026)
by: Lin, Yihan, et al.
Published: (2026)
Benchmarking Large Language Models for Geolocating Colonial Virginia Land Grants
by: Mioduski, Ryan
Published: (2025)
by: Mioduski, Ryan
Published: (2025)
Truth Without Comprehension: A BlueSky Agenda for Steering the Fourth Mathematical Crisis
by: Yu, Runlong, et al.
Published: (2025)
by: Yu, Runlong, et al.
Published: (2025)
GeoRouter: Dynamic Paradigm Routing for Worldwide Image Geolocalization
by: Jia, Pengyue, et al.
Published: (2026)
by: Jia, Pengyue, et al.
Published: (2026)
Where on Earth? A Vision-Language Benchmark for Probing Model Geolocation Skills Across Scales
by: Qian, Zhaofang, et al.
Published: (2025)
by: Qian, Zhaofang, et al.
Published: (2025)
Evaluating the Diagnostic Classification Ability of Multimodal Large Language Models: Insights from the Osteoarthritis Initiative
by: Wang, Li, et al.
Published: (2026)
by: Wang, Li, et al.
Published: (2026)
Q-Bench-Portrait: Benchmarking Multimodal Large Language Models on Portrait Image Quality Perception
by: Wu, Sijing, et al.
Published: (2026)
by: Wu, Sijing, et al.
Published: (2026)
Medical Large Vision Language Models with Multi-Image Visual Ability
by: Yang, Xikai, et al.
Published: (2025)
by: Yang, Xikai, et al.
Published: (2025)
From Waveforms to Pixels: A Survey on Audio-Visual Segmentation
by: Li, Jia, et al.
Published: (2025)
by: Li, Jia, et al.
Published: (2025)
Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens
by: Chen, Feng, et al.
Published: (2024)
by: Chen, Feng, et al.
Published: (2024)
ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models
by: Wang, Ziyue, et al.
Published: (2024)
by: Wang, Ziyue, et al.
Published: (2024)
VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models
by: Li, Chenglin, et al.
Published: (2024)
by: Li, Chenglin, et al.
Published: (2024)
Plug to Place: Indoor Multimedia Geolocation from Electrical Sockets for Digital Investigation
by: Aftab, Kanwal, et al.
Published: (2025)
by: Aftab, Kanwal, et al.
Published: (2025)
GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing
by: Ou, Ruizhe, et al.
Published: (2025)
by: Ou, Ruizhe, et al.
Published: (2025)
NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes
by: Fan, Lizhou, et al.
Published: (2023)
by: Fan, Lizhou, et al.
Published: (2023)
Towards a Multimodal Large Language Model with Pixel-Level Insight for Biomedicine
by: Huang, Xiaoshuang, et al.
Published: (2024)
by: Huang, Xiaoshuang, et al.
Published: (2024)
Evaluation of Geolocation Capabilities of Multimodal Large Language Models and Analysis of Associated Privacy Risks
by: Zhang, Xian, et al.
Published: (2025)
by: Zhang, Xian, et al.
Published: (2025)
CoSpace: Benchmarking Continuous Space Perception Ability for Vision-Language Models
by: Zhu, Yiqi, et al.
Published: (2025)
by: Zhu, Yiqi, et al.
Published: (2025)
From Pixel to Mask: A Survey of Out-of-Distribution Segmentation
by: Zhao, Wenjie, et al.
Published: (2025)
by: Zhao, Wenjie, et al.
Published: (2025)
DT-NeRF: A Diffusion and Transformer-Based Optimization Approach for Neural Radiance Fields in 3D Reconstruction
by: Liu, Bo, et al.
Published: (2025)
by: Liu, Bo, et al.
Published: (2025)
From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language Models
by: Shang, Yuying, et al.
Published: (2024)
by: Shang, Yuying, et al.
Published: (2024)
LSDTs: LLM-Augmented Semantic Digital Twins for Adaptive Knowledge-Intensive Infrastructure Planning
by: Li, Naiyi, et al.
Published: (2025)
by: Li, Naiyi, et al.
Published: (2025)
PiTe: Pixel-Temporal Alignment for Large Video-Language Model
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models
by: Kang, Zhaolu, et al.
Published: (2025)
by: Kang, Zhaolu, et al.
Published: (2025)
From Image- to Pixel-level: Label-efficient Hyperspectral Image Reconstruction
by: Leng, Yihong, et al.
Published: (2025)
by: Leng, Yihong, et al.
Published: (2025)
PIGEON: Predicting Image Geolocations
by: Haas, Lukas, et al.
Published: (2023)
by: Haas, Lukas, et al.
Published: (2023)
Similar Items
-
Empowering LLM Agents with Geospatial Awareness: Toward Grounded Reasoning for Wildfire Response
by: Chen, Yiheng, et al.
Published: (2025) -
LLMGeo: Benchmarking Large Language Models on Image Geolocation In-the-wild
by: Wang, Zhiqiang, et al.
Published: (2024) -
When Earth Foundation Models Meet Diffusion: An Application to Land Surface Temperature Super-Resolution
by: Chen, Yiheng, et al.
Published: (2026) -
Image-Based Geolocation Using Large Vision-Language Models
by: Liu, Yi, et al.
Published: (2024) -
Learning to Wander: Improving the Global Image Geolocation Ability of LMMs via Actionable Reasoning
by: Zheng, Yushuo, et al.
Published: (2026)