LLMGeo: Benchmarking Large Language Models on Image Geolocation In-the-wild

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zhiqiang, Xu, Dejia, Khan, Rana Muhammad Shahroz, Lin, Yanbin, Fan, Zhiwen, Zhu, Xingquan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911895277010944
author Wang, Zhiqiang
Xu, Dejia
Khan, Rana Muhammad Shahroz
Lin, Yanbin
Fan, Zhiwen
Zhu, Xingquan
author_facet Wang, Zhiqiang
Xu, Dejia
Khan, Rana Muhammad Shahroz
Lin, Yanbin
Fan, Zhiwen
Zhu, Xingquan
contents Image geolocation is a critical task in various image-understanding applications. However, existing methods often fail when analyzing challenging, in-the-wild images. Inspired by the exceptional background knowledge of multimodal language models, we systematically evaluate their geolocation capabilities using a novel image dataset and a comprehensive evaluation framework. We first collect images from various countries via Google Street View. Then, we conduct training-free and training-based evaluations on closed-source and open-source multi-modal language models. we conduct both training-free and training-based evaluations on closed-source and open-source multimodal language models. Our findings indicate that closed-source models demonstrate superior geolocation abilities, while open-source models can achieve comparable performance through fine-tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2405_20363
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LLMGeo: Benchmarking Large Language Models on Image Geolocation In-the-wild
Wang, Zhiqiang
Xu, Dejia
Khan, Rana Muhammad Shahroz
Lin, Yanbin
Fan, Zhiwen
Zhu, Xingquan
Computer Vision and Pattern Recognition
Image geolocation is a critical task in various image-understanding applications. However, existing methods often fail when analyzing challenging, in-the-wild images. Inspired by the exceptional background knowledge of multimodal language models, we systematically evaluate their geolocation capabilities using a novel image dataset and a comprehensive evaluation framework. We first collect images from various countries via Google Street View. Then, we conduct training-free and training-based evaluations on closed-source and open-source multi-modal language models. we conduct both training-free and training-based evaluations on closed-source and open-source multimodal language models. Our findings indicate that closed-source models demonstrate superior geolocation abilities, while open-source models can achieve comparable performance through fine-tuning.
title LLMGeo: Benchmarking Large Language Models on Image Geolocation In-the-wild
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.20363