Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bicakci, Yunus Serhat, Shingleton, Joseph, Basiri, Anahid |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multimodal Retrieval-Augmented Generation with Large Language Models for Medical VQA
von: Karim, A H M Rezaul, et al.
Veröffentlicht: (2025)
von: Karim, A H M Rezaul, et al.
Veröffentlicht: (2025)
Zero-shot Vision-Language Reranking for Cross-View Geolocalization
von: Erzurumlu, Yunus Talha, et al.
Veröffentlicht: (2026)
von: Erzurumlu, Yunus Talha, et al.
Veröffentlicht: (2026)
UlcerGPT: A Multimodal Approach Leveraging Large Language and Vision Models for Diabetic Foot Ulcer Image Transcription
von: Basiri, Reza, et al.
Veröffentlicht: (2024)
von: Basiri, Reza, et al.
Veröffentlicht: (2024)
Img2Loc: Revisiting Image Geolocalization using Multi-modality Foundation Models and Image-based Retrieval-Augmented Generation
von: Zhou, Zhongliang, et al.
Veröffentlicht: (2024)
von: Zhou, Zhongliang, et al.
Veröffentlicht: (2024)
OpenStreetView-5M: The Many Roads to Global Visual Geolocation
von: Astruc, Guillaume, et al.
Veröffentlicht: (2024)
von: Astruc, Guillaume, et al.
Veröffentlicht: (2024)
Skill-Conditioned Visual Geolocation for Vision-Language Models
von: Yang, Chenjie, et al.
Veröffentlicht: (2026)
von: Yang, Chenjie, et al.
Veröffentlicht: (2026)
G3: An Effective and Adaptive Framework for Worldwide Geolocalization Using Large Multi-Modality Models
von: Jia, Pengyue, et al.
Veröffentlicht: (2024)
von: Jia, Pengyue, et al.
Veröffentlicht: (2024)
RAP: Retrieval-Augmented Personalization for Multimodal Large Language Models
von: Hao, Haoran, et al.
Veröffentlicht: (2024)
von: Hao, Haoran, et al.
Veröffentlicht: (2024)
Cross-View Geolocalization and Disaster Mapping with Street-View and VHR Satellite Imagery: A Case Study of Hurricane IAN
von: Li, Hao, et al.
Veröffentlicht: (2024)
von: Li, Hao, et al.
Veröffentlicht: (2024)
Towards a Multimodal Large Language Model with Pixel-Level Insight for Biomedicine
von: Huang, Xiaoshuang, et al.
Veröffentlicht: (2024)
von: Huang, Xiaoshuang, et al.
Veröffentlicht: (2024)
On the Out-Of-Distribution Generalization of Multimodal Large Language Models
von: Zhang, Xingxuan, et al.
Veröffentlicht: (2024)
von: Zhang, Xingxuan, et al.
Veröffentlicht: (2024)
Alleviating Hallucination in Large Vision-Language Models with Active Retrieval Augmentation
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
Thinking with Map: Reinforced Parallel Map-Augmented Agent for Geolocalization
von: Ji, Yuxiang, et al.
Veröffentlicht: (2026)
von: Ji, Yuxiang, et al.
Veröffentlicht: (2026)
Pixel-Grounded Retrieval for Knowledgeable Large Multimodal Models
von: Kim, Jeonghwan, et al.
Veröffentlicht: (2026)
von: Kim, Jeonghwan, et al.
Veröffentlicht: (2026)
UniFashion: A Unified Vision-Language Model for Multimodal Fashion Retrieval and Generation
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2024)
Examining the Commitments and Difficulties Inherent in Multimodal Foundation Models for Street View Imagery
von: Yang, Zhenyuan, et al.
Veröffentlicht: (2024)
von: Yang, Zhenyuan, et al.
Veröffentlicht: (2024)
Retrieval-Augmented Dynamic Prompt Tuning for Incomplete Multimodal Learning
von: Lang, Jian, et al.
Veröffentlicht: (2025)
von: Lang, Jian, et al.
Veröffentlicht: (2025)
GeoShield: Safeguarding Geolocation Privacy from Vision-Language Models via Adversarial Perturbations
von: Liu, Xinwei, et al.
Veröffentlicht: (2025)
von: Liu, Xinwei, et al.
Veröffentlicht: (2025)
PARTONOMY: Large Multimodal Models with Part-Level Visual Understanding
von: Blume, Ansel, et al.
Veröffentlicht: (2025)
von: Blume, Ansel, et al.
Veröffentlicht: (2025)
BuildingView: Constructing Urban Building Exteriors Databases with Street View Imagery and Multimodal Large Language Mode
von: Li, Zongrong, et al.
Veröffentlicht: (2024)
von: Li, Zongrong, et al.
Veröffentlicht: (2024)
Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation
von: Sanguigni, Fulvio, et al.
Veröffentlicht: (2025)
von: Sanguigni, Fulvio, et al.
Veröffentlicht: (2025)
mKG-RAG: Leveraging Multimodal Knowledge Graphs in Retrieval-Augmented Generation for Knowledge-intensive VQA
von: Yuan, Xu, et al.
Veröffentlicht: (2025)
von: Yuan, Xu, et al.
Veröffentlicht: (2025)
MMS-VPR: Multimodal Street-Level Visual Place Recognition Dataset and Benchmark
von: Ou, Yiwei, et al.
Veröffentlicht: (2025)
von: Ou, Yiwei, et al.
Veröffentlicht: (2025)
MMR-AD: A Large-Scale Multimodal Dataset for Benchmarking General Anomaly Detection with Multimodal Large Language Models
von: Yao, Xincheng, et al.
Veröffentlicht: (2026)
von: Yao, Xincheng, et al.
Veröffentlicht: (2026)
Remote Sensing Retrieval-Augmented Generation: Bridging Remote Sensing Imagery and Comprehensive Knowledge with a Multi-Modal Dataset and Retrieval-Augmented Generation Model
von: Wen, Congcong, et al.
Veröffentlicht: (2025)
von: Wen, Congcong, et al.
Veröffentlicht: (2025)
MRAG-Bench: Vision-Centric Evaluation for Retrieval-Augmented Multimodal Models
von: Hu, Wenbo, et al.
Veröffentlicht: (2024)
von: Hu, Wenbo, et al.
Veröffentlicht: (2024)
Robust Adaptation of Large Multimodal Models for Retrieval Augmented Hateful Meme Detection
von: Mei, Jingbiao, et al.
Veröffentlicht: (2025)
von: Mei, Jingbiao, et al.
Veröffentlicht: (2025)
Genixer: Empowering Multimodal Large Language Models as a Powerful Data Generator
von: Zhao, Henry Hengyuan, et al.
Veröffentlicht: (2023)
von: Zhao, Henry Hengyuan, et al.
Veröffentlicht: (2023)
Knowledge Completes the Vision: A Multimodal Entity-aware Retrieval-Augmented Generation Framework for News Image Captioning
von: You, Xiaoxing, et al.
Veröffentlicht: (2025)
von: You, Xiaoxing, et al.
Veröffentlicht: (2025)
VLADriver-RAG: Retrieval-Augmented Vision-Language-Action Models for Autonomous Driving
von: Zhao, Rui, et al.
Veröffentlicht: (2026)
von: Zhao, Rui, et al.
Veröffentlicht: (2026)
Leveraging Retrieval Augment Approach for Multimodal Emotion Recognition Under Missing Modalities
von: Fan, Qi, et al.
Veröffentlicht: (2024)
von: Fan, Qi, et al.
Veröffentlicht: (2024)
Sat3DGen: Comprehensive Street-Level 3D Scene Generation from Single Satellite Image
von: Qian, Ming, et al.
Veröffentlicht: (2026)
von: Qian, Ming, et al.
Veröffentlicht: (2026)
Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs
von: Caffagni, Davide, et al.
Veröffentlicht: (2024)
von: Caffagni, Davide, et al.
Veröffentlicht: (2024)
HV-Attack: Hierarchical Visual Attack for Multimodal Retrieval Augmented Generation
von: Luo, Linyin, et al.
Veröffentlicht: (2025)
von: Luo, Linyin, et al.
Veröffentlicht: (2025)
Mesh RAG: Retrieval Augmentation for Autoregressive Mesh Generation
von: Sun, Xiatao, et al.
Veröffentlicht: (2025)
von: Sun, Xiatao, et al.
Veröffentlicht: (2025)
Retrieval-Augmented Anatomical Guidance for Text-to-CT Generation
von: Molino, Daniele, et al.
Veröffentlicht: (2026)
von: Molino, Daniele, et al.
Veröffentlicht: (2026)
ChartInsights: Evaluating Multimodal Large Language Models for Low-Level Chart Question Answering
von: Wu, Yifan, et al.
Veröffentlicht: (2024)
von: Wu, Yifan, et al.
Veröffentlicht: (2024)
Speculative Decoding Reimagined for Multimodal Large Language Models
von: Lin, Luxi, et al.
Veröffentlicht: (2025)
von: Lin, Luxi, et al.
Veröffentlicht: (2025)
Efficient Multimodal Large Language Models: A Survey
von: Jin, Yizhang, et al.
Veröffentlicht: (2024)
von: Jin, Yizhang, et al.
Veröffentlicht: (2024)
Toward Cognitive Supersensing in Multimodal Large Language Model
von: Li, Boyi, et al.
Veröffentlicht: (2026)
von: Li, Boyi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Multimodal Retrieval-Augmented Generation with Large Language Models for Medical VQA
von: Karim, A H M Rezaul, et al.
Veröffentlicht: (2025) -
Zero-shot Vision-Language Reranking for Cross-View Geolocalization
von: Erzurumlu, Yunus Talha, et al.
Veröffentlicht: (2026) -
UlcerGPT: A Multimodal Approach Leveraging Large Language and Vision Models for Diabetic Foot Ulcer Image Transcription
von: Basiri, Reza, et al.
Veröffentlicht: (2024) -
Img2Loc: Revisiting Image Geolocalization using Multi-modality Foundation Models and Image-based Retrieval-Augmented Generation
von: Zhou, Zhongliang, et al.
Veröffentlicht: (2024) -
OpenStreetView-5M: The Many Roads to Global Visual Geolocation
von: Astruc, Guillaume, et al.
Veröffentlicht: (2024)