Zero-shot Vision-Language Reranking for Cross-View Geolocalization
Fuente:
arXiv
Saved in:
| Main Authors: | Erzurumlu, Yunus Talha, Anderson, John E., Shuart, William J., Toth, Charles, Yilmaz, Alper |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Just Zoom In: Cross-View Geo-Localization via Autoregressive Zooming
by: Erzurumlu, Yunus Talha, et al.
Published: (2026)
by: Erzurumlu, Yunus Talha, et al.
Published: (2026)
BetterScene: 3D Scene Synthesis with Representation-Aligned Generative Model
by: Han, Yuci, et al.
Published: (2026)
by: Han, Yuci, et al.
Published: (2026)
MobileUNETR: A Lightweight End-To-End Hybrid Vision Transformer For Efficient Medical Image Segmentation
by: Perera, Shehan, et al.
Published: (2024)
by: Perera, Shehan, et al.
Published: (2024)
Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation
by: Bicakci, Yunus Serhat, et al.
Published: (2025)
by: Bicakci, Yunus Serhat, et al.
Published: (2025)
Skill-Conditioned Visual Geolocation for Vision-Language Models
by: Yang, Chenjie, et al.
Published: (2026)
by: Yang, Chenjie, et al.
Published: (2026)
Gate-and-Merge: Zero-shot Compositional Personalization of Vision Language Models
by: Ding, Guodong, et al.
Published: (2026)
by: Ding, Guodong, et al.
Published: (2026)
Cross-View Geolocalization and Disaster Mapping with Street-View and VHR Satellite Imagery: A Case Study of Hurricane IAN
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
OpenStreetView-5M: The Many Roads to Global Visual Geolocation
by: Astruc, Guillaume, et al.
Published: (2024)
by: Astruc, Guillaume, et al.
Published: (2024)
P2DNav: Panorama-to-Downview Reasoning for Zero-shot Vision-and-Language Navigation
by: Sheng, Kai, et al.
Published: (2026)
by: Sheng, Kai, et al.
Published: (2026)
Zero-shot detection of buildings in mobile LiDAR using Language Vision Model
by: Goo, June Moh, et al.
Published: (2024)
by: Goo, June Moh, et al.
Published: (2024)
Understanding Graphical Perception in Data Visualization through Zero-shot Prompting of Vision-Language Models
by: Guo, Grace, et al.
Published: (2024)
by: Guo, Grace, et al.
Published: (2024)
GeoShield: Safeguarding Geolocation Privacy from Vision-Language Models via Adversarial Perturbations
by: Liu, Xinwei, et al.
Published: (2025)
by: Liu, Xinwei, et al.
Published: (2025)
A Vision-Language Foundation Model for Zero-shot Clinical Collaboration and Automated Concept Discovery in Dermatology
by: Yan, Siyuan, et al.
Published: (2026)
by: Yan, Siyuan, et al.
Published: (2026)
Lightweight Road Environment Segmentation using Vector Quantization
by: Kwag, Jiyong, et al.
Published: (2025)
by: Kwag, Jiyong, et al.
Published: (2025)
Deep learning-based automated damage detection in concrete structures using images from earthquake events
by: Turer, Abdullah, et al.
Published: (2025)
by: Turer, Abdullah, et al.
Published: (2025)
Multivariate Gaussian Representation Learning for Medical Action Evaluation
by: Yang, Luming, et al.
Published: (2025)
by: Yang, Luming, et al.
Published: (2025)
Hierarchically Robust Zero-shot Vision-language Models
by: Dong, Junhao, et al.
Published: (2026)
by: Dong, Junhao, et al.
Published: (2026)
UAS Visual Navigation in Large and Unseen Environments via a Meta Agent
by: Han, Yuci, et al.
Published: (2025)
by: Han, Yuci, et al.
Published: (2025)
GRAZE: Grounded Refinement and Motion-Aware Zero-Shot Event Localization
by: Zaidi, Syed Ahsan Masud, et al.
Published: (2026)
by: Zaidi, Syed Ahsan Masud, et al.
Published: (2026)
Computer Vision for Multimedia Geolocation in Human Trafficking Investigation: A Systematic Literature Review
by: Bamigbade, Opeyemi, et al.
Published: (2024)
by: Bamigbade, Opeyemi, et al.
Published: (2024)
Zoom-shot: Fast and Efficient Unsupervised Zero-Shot Transfer of CLIP to Vision Encoders with Multimodal Loss
by: Shipard, Jordan, et al.
Published: (2024)
by: Shipard, Jordan, et al.
Published: (2024)
Towards Generative Location Awareness for Disaster Response: A Probabilistic Cross-view Geolocalization Approach
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
Visual Space Optimization for Zero-shot Learning
by: Wang, Xinsheng, et al.
Published: (2019)
by: Wang, Xinsheng, et al.
Published: (2019)
Fine-gained Zero-shot Video Sampling
by: Chen, Dengsheng, et al.
Published: (2024)
by: Chen, Dengsheng, et al.
Published: (2024)
HecVL: Hierarchical Video-Language Pretraining for Zero-shot Surgical Phase Recognition
by: Yuan, Kun, et al.
Published: (2024)
by: Yuan, Kun, et al.
Published: (2024)
R3G: A Reasoning--Retrieval--Reranking Framework for Vision-Centric Answer Generation
by: Chen, Zhuohong, et al.
Published: (2026)
by: Chen, Zhuohong, et al.
Published: (2026)
TopV-Nav: Unlocking the Top-View Spatial Reasoning Potential of MLLM for Zero-shot Object Navigation
by: Zhong, Linqing, et al.
Published: (2024)
by: Zhong, Linqing, et al.
Published: (2024)
Zero-shot World Models Are Developmentally Efficient Learners
by: Aw, Khai Loong, et al.
Published: (2026)
by: Aw, Khai Loong, et al.
Published: (2026)
CrossVL: Complexity-Aware Feature Routing and Paired Curriculum for Cross-View Vision-Language Detection
by: Liu, Zhipeng, et al.
Published: (2026)
by: Liu, Zhipeng, et al.
Published: (2026)
Revealing Multi-View Hallucination in Large Vision-Language Models
by: Park, Wooje, et al.
Published: (2026)
by: Park, Wooje, et al.
Published: (2026)
OAD-Promoter: Enhancing Zero-shot VQA using Large Language Models with Object Attribute Description
by: Xu, Quanxing, et al.
Published: (2025)
by: Xu, Quanxing, et al.
Published: (2025)
AIR: Zero-shot Generative Model Adaptation with Iterative Refinement
by: Liu, Guimeng, et al.
Published: (2025)
by: Liu, Guimeng, et al.
Published: (2025)
Conceptrol: Concept Control of Zero-shot Personalized Image Generation
by: He, Qiyuan, et al.
Published: (2025)
by: He, Qiyuan, et al.
Published: (2025)
Zero-shot High-fidelity and Pose-controllable Character Animation
by: Zhu, Bingwen, et al.
Published: (2024)
by: Zhu, Bingwen, et al.
Published: (2024)
InstantID: Zero-shot Identity-Preserving Generation in Seconds
by: Wang, Qixun, et al.
Published: (2024)
by: Wang, Qixun, et al.
Published: (2024)
Anomagic: Crossmodal Prompt-driven Zero-shot Anomaly Generation
by: Jiang, Yuxin, et al.
Published: (2025)
by: Jiang, Yuxin, et al.
Published: (2025)
Prompt-based Visual Alignment for Zero-shot Policy Transfer
by: Gao, Haihan, et al.
Published: (2024)
by: Gao, Haihan, et al.
Published: (2024)
Cross-domain Multi-step Thinking: Zero-shot Fine-grained Traffic Sign Recognition in the Wild
by: Gan, Yaozong, et al.
Published: (2024)
by: Gan, Yaozong, et al.
Published: (2024)
Toward an Artificial General Teacher: Procedural Geometry Data Generation and Visual Grounding with Vision-Language Models
by: Nguyen-Truong, Hai, et al.
Published: (2026)
by: Nguyen-Truong, Hai, et al.
Published: (2026)
Can Medical Vision-Language Pre-training Succeed with Purely Synthetic Data?
by: Liu, Che, et al.
Published: (2024)
by: Liu, Che, et al.
Published: (2024)
Similar Items
-
Just Zoom In: Cross-View Geo-Localization via Autoregressive Zooming
by: Erzurumlu, Yunus Talha, et al.
Published: (2026) -
BetterScene: 3D Scene Synthesis with Representation-Aligned Generative Model
by: Han, Yuci, et al.
Published: (2026) -
MobileUNETR: A Lightweight End-To-End Hybrid Vision Transformer For Efficient Medical Image Segmentation
by: Perera, Shehan, et al.
Published: (2024) -
Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation
by: Bicakci, Yunus Serhat, et al.
Published: (2025) -
Skill-Conditioned Visual Geolocation for Vision-Language Models
by: Yang, Chenjie, et al.
Published: (2026)