Img2Loc: Revisiting Image Geolocalization using Multi-modality Foundation Models and Image-based Retrieval-Augmented Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Zhongliang, Zhang, Jielu, Guan, Zihan, Hu, Mengxuan, Lao, Ni, Mu, Lan, Li, Sheng, Mai, Gengchen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Text2Seg: Remote Sensing Image Semantic Segmentation via Text-Guided Visual Foundation Models
von: Zhang, Jielu, et al.
Veröffentlicht: (2023)
von: Zhang, Jielu, et al.
Veröffentlicht: (2023)
LocDiff: Identifying Locations on Earth by Diffusing in the Hilbert Space
von: Wang, Zhangyu, et al.
Veröffentlicht: (2025)
von: Wang, Zhangyu, et al.
Veröffentlicht: (2025)
Enhanced Diagnostic Performance via Large-Resolution Inference Optimization for Pathology Foundation Models
von: Hu, Mengxuan, et al.
Veröffentlicht: (2026)
von: Hu, Mengxuan, et al.
Veröffentlicht: (2026)
BalancEdit: Dynamically Balancing the Generality-Locality Trade-off in Multi-modal Model Editing
von: Guo, Dongliang, et al.
Veröffentlicht: (2025)
von: Guo, Dongliang, et al.
Veröffentlicht: (2025)
MC-GTA: Metric-Constrained Model-Based Clustering using Goodness-of-fit Tests with Autocorrelations
von: Wang, Zhangyu, et al.
Veröffentlicht: (2024)
von: Wang, Zhangyu, et al.
Veröffentlicht: (2024)
No Free Lunch: Retrieval-Augmented Generation Undermines Fairness in LLMs, Even for Vigilant Users
von: Hu, Mengxuan, et al.
Veröffentlicht: (2024)
von: Hu, Mengxuan, et al.
Veröffentlicht: (2024)
GAIR: Location-Aware Self-Supervised Contrastive Pre-Training with Geo-Aligned Implicit Representations
von: Liu, Zeping, et al.
Veröffentlicht: (2025)
von: Liu, Zeping, et al.
Veröffentlicht: (2025)
TorchSpatial: A Location Encoding Framework and Benchmark for Spatial Representation Learning
von: Wu, Nemin, et al.
Veröffentlicht: (2024)
von: Wu, Nemin, et al.
Veröffentlicht: (2024)
CurriculumLoc: Enhancing Cross-Domain Geolocalization through Multi-Stage Refinement
von: Hu, Boni, et al.
Veröffentlicht: (2023)
von: Hu, Boni, et al.
Veröffentlicht: (2023)
UFID: A Unified Framework for Input-level Backdoor Detection on Diffusion Models
von: Guan, Zihan, et al.
Veröffentlicht: (2024)
von: Guan, Zihan, et al.
Veröffentlicht: (2024)
ImLoc: Revisiting Visual Localization with Image-based Representation
von: Jiang, Xudong, et al.
Veröffentlicht: (2026)
von: Jiang, Xudong, et al.
Veröffentlicht: (2026)
Benign Samples Matter! Fine-tuning On Outlier Benign Samples Severely Breaks Safety
von: Guan, Zihan, et al.
Veröffentlicht: (2025)
von: Guan, Zihan, et al.
Veröffentlicht: (2025)
GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models
von: Yi, Qiang, et al.
Veröffentlicht: (2025)
von: Yi, Qiang, et al.
Veröffentlicht: (2025)
Provably Secure Retrieval-Augmented Generation
von: Zhou, Pengcheng, et al.
Veröffentlicht: (2025)
von: Zhou, Pengcheng, et al.
Veröffentlicht: (2025)
HierLoc: Hyperbolic Entity Embeddings for Hierarchical Visual Geolocation
von: Gadi, Hari Krishna, et al.
Veröffentlicht: (2026)
von: Gadi, Hari Krishna, et al.
Veröffentlicht: (2026)
Multi-modal Retrieval Augmented Multi-modal Generation: Datasets, Evaluation Metrics and Strong Baselines
von: Ma, Zi-Ao, et al.
Veröffentlicht: (2024)
von: Ma, Zi-Ao, et al.
Veröffentlicht: (2024)
ProxyImg: Towards Highly-Controllable Image Representation via Hierarchical Disentangled Proxy Embedding
von: Chen, Ye, et al.
Veröffentlicht: (2026)
von: Chen, Ye, et al.
Veröffentlicht: (2026)
ImgTrojan: Jailbreaking Vision-Language Models with ONE Image
von: Tao, Xijia, et al.
Veröffentlicht: (2024)
von: Tao, Xijia, et al.
Veröffentlicht: (2024)
ImgEdit: A Unified Image Editing Dataset and Benchmark
von: Ye, Yang, et al.
Veröffentlicht: (2025)
von: Ye, Yang, et al.
Veröffentlicht: (2025)
Spatial-RAG: Spatial Retrieval Augmented Generation for Real-World Geospatial Reasoning Questions
von: Yu, Dazhou, et al.
Veröffentlicht: (2025)
von: Yu, Dazhou, et al.
Veröffentlicht: (2025)
ManuRAG: Multi-modal Retrieval Augmented Generation for Manufacturing Question Answering (Early Version)
von: Li, Yunqing, et al.
Veröffentlicht: (2026)
von: Li, Yunqing, et al.
Veröffentlicht: (2026)
Feature-Augmented Deep Networks for Multiscale Building Segmentation in High-Resolution UAV and Satellite Imagery
von: Maniyar, Chintan B., et al.
Veröffentlicht: (2025)
von: Maniyar, Chintan B., et al.
Veröffentlicht: (2025)
Image Quality Assessment: Exploring Regional Heterogeneity via Response of Adaptive Multiple Quality Factors in Dictionary Space
von: Lan, Xuting, et al.
Veröffentlicht: (2024)
von: Lan, Xuting, et al.
Veröffentlicht: (2024)
Backdoor in Seconds: Unlocking Vulnerabilities in Large Pre-trained Models via Model Editing
von: Guo, Dongliang, et al.
Veröffentlicht: (2024)
von: Guo, Dongliang, et al.
Veröffentlicht: (2024)
AesExpert: Towards Multi-modality Foundation Model for Image Aesthetics Perception
von: Huang, Yipo, et al.
Veröffentlicht: (2024)
von: Huang, Yipo, et al.
Veröffentlicht: (2024)
GeoSearch: Augmenting Worldwide Geolocalization with Web-Scale Reverse Image Search and Image Matching
von: Le-Duc, Tung-Duong, et al.
Veröffentlicht: (2026)
von: Le-Duc, Tung-Duong, et al.
Veröffentlicht: (2026)
TRAJGANR: Trajectory-Centric Urban Multimodal Learning via Geospatially Aligned Neural Representations
von: Siampou, Maria Despoina, et al.
Veröffentlicht: (2026)
von: Siampou, Maria Despoina, et al.
Veröffentlicht: (2026)
PIGEON: Predicting Image Geolocations
von: Haas, Lukas, et al.
Veröffentlicht: (2023)
von: Haas, Lukas, et al.
Veröffentlicht: (2023)
Image Quality Assessment: Exploring Quality Awareness via Memory-driven Distortion Patterns Matching
von: Lan, Xuting, et al.
Veröffentlicht: (2026)
von: Lan, Xuting, et al.
Veröffentlicht: (2026)
Img2CADSeq: Image-to-CAD Generation via Sequence-Based Diffusion
von: Tan, Shiyu, et al.
Veröffentlicht: (2026)
von: Tan, Shiyu, et al.
Veröffentlicht: (2026)
Cross-View Geolocalization and Disaster Mapping with Street-View and VHR Satellite Imagery: A Case Study of Hurricane IAN
von: Li, Hao, et al.
Veröffentlicht: (2024)
von: Li, Hao, et al.
Veröffentlicht: (2024)
Cross-modal RAG: Sub-dimensional Text-to-Image Retrieval-Augmented Generation
von: Zhu, Mengdan, et al.
Veröffentlicht: (2025)
von: Zhu, Mengdan, et al.
Veröffentlicht: (2025)
Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models
von: Lai, Zhengfeng, et al.
Veröffentlicht: (2024)
von: Lai, Zhengfeng, et al.
Veröffentlicht: (2024)
Retrieval Augmented Image Harmonization
von: Wang, Haolin, et al.
Veröffentlicht: (2024)
von: Wang, Haolin, et al.
Veröffentlicht: (2024)
Quantitative Phase Imaging for Meta‐Lenses by Phase Retrieval
von: Jialuo Cheng, et al.
Veröffentlicht: (2025)
von: Jialuo Cheng, et al.
Veröffentlicht: (2025)
Neural Architecture Search generated Phase Retrieval Net for Real-time Off-axis Quantitative Phase Imaging
von: Shu, Xin, et al.
Veröffentlicht: (2022)
von: Shu, Xin, et al.
Veröffentlicht: (2022)
3DProxyImg: Controllable 3D-Aware Animation Synthesis from Single Image via 2D-3D Aligned Proxy Embedding
von: Zhu, Yupeng, et al.
Veröffentlicht: (2025)
von: Zhu, Yupeng, et al.
Veröffentlicht: (2025)
Multi-modal Reference Learning for Fine-grained Text-to-Image Retrieval
von: Ma, Zehong, et al.
Veröffentlicht: (2025)
von: Ma, Zehong, et al.
Veröffentlicht: (2025)
6Img-to-3D: Few-Image Large-Scale Outdoor Driving Scene Reconstruction
von: Gieruc, Théo, et al.
Veröffentlicht: (2024)
von: Gieruc, Théo, et al.
Veröffentlicht: (2024)
GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation
von: Yan, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Yan, Zhiyuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Text2Seg: Remote Sensing Image Semantic Segmentation via Text-Guided Visual Foundation Models
von: Zhang, Jielu, et al.
Veröffentlicht: (2023) -
LocDiff: Identifying Locations on Earth by Diffusing in the Hilbert Space
von: Wang, Zhangyu, et al.
Veröffentlicht: (2025) -
Enhanced Diagnostic Performance via Large-Resolution Inference Optimization for Pathology Foundation Models
von: Hu, Mengxuan, et al.
Veröffentlicht: (2026) -
BalancEdit: Dynamically Balancing the Generality-Locality Trade-off in Multi-modal Model Editing
von: Guo, Dongliang, et al.
Veröffentlicht: (2025) -
MC-GTA: Metric-Constrained Model-Based Clustering using Goodness-of-fit Tests with Autocorrelations
von: Wang, Zhangyu, et al.
Veröffentlicht: (2024)