Where on Earth? A Vision-Language Benchmark for Probing Model Geolocation Skills Across Scales
Fuente:
arXiv
Saved in:
| Main Authors: | Qian, Zhaofang, Chen, Hardy, Wang, Zeyu, Zhang, Li, Wang, Zijun, Huang, Xiaoke, Liu, Hui, Tang, Xianfeng, Zheng, Zeyu, Tu, Haoqin, Xie, Cihang, Zhou, Yuyin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
by: Chen, Hardy, et al.
Published: (2025)
by: Chen, Hardy, et al.
Published: (2025)
ViLBench: A Suite for Vision-Language Process Reward Modeling
by: Tu, Haoqin, et al.
Published: (2025)
by: Tu, Haoqin, et al.
Published: (2025)
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
by: Wu, Juncheng, et al.
Published: (2026)
by: Wu, Juncheng, et al.
Published: (2026)
ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoning
by: Wu, Juncheng, et al.
Published: (2026)
by: Wu, Juncheng, et al.
Published: (2026)
Kestrel: Grounding Self-Refinement for LVLM Hallucination Mitigation
by: Mao, Jiawei, et al.
Published: (2026)
by: Mao, Jiawei, et al.
Published: (2026)
Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains
by: Wu, Juncheng, et al.
Published: (2025)
by: Wu, Juncheng, et al.
Published: (2025)
Sculpting Holistic 3D Representation in Contrastive Language-Image-3D Pre-training
by: Gao, Yipeng, et al.
Published: (2023)
by: Gao, Yipeng, et al.
Published: (2023)
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning
by: Liu, Yanqing, et al.
Published: (2025)
by: Liu, Yanqing, et al.
Published: (2025)
m1: Unleash the Potential of Test-Time Scaling for Medical Reasoning with Large Language Models
by: Huang, Xiaoke, et al.
Published: (2025)
by: Huang, Xiaoke, et al.
Published: (2025)
Your Agent, Their Asset: A Real-World Safety Analysis of OpenClaw
by: Wang, Zijun, et al.
Published: (2026)
by: Wang, Zijun, et al.
Published: (2026)
AttnGCG: Enhancing Jailbreaking Attacks on LLMs with Attention Manipulation
by: Wang, Zijun, et al.
Published: (2024)
by: Wang, Zijun, et al.
Published: (2024)
Story-Iter: A Training-free Iterative Paradigm for Long Story Visualization
by: Mao, Jiawei, et al.
Published: (2024)
by: Mao, Jiawei, et al.
Published: (2024)
Synthesizing High-Quality Visual Question Answering from Medical Documents with Generator-Verifier LMMs
by: Huang, Xiaoke, et al.
Published: (2025)
by: Huang, Xiaoke, et al.
Published: (2025)
What If We Recaption Billions of Web Images with LLaMA-3?
by: Li, Xianhang, et al.
Published: (2024)
by: Li, Xianhang, et al.
Published: (2024)
Revisiting Adversarial Training at Scale
by: Wang, Zeyu, et al.
Published: (2024)
by: Wang, Zeyu, et al.
Published: (2024)
OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning
by: Li, Xianhang, et al.
Published: (2025)
by: Li, Xianhang, et al.
Published: (2025)
AHELM: A Holistic Evaluation of Audio-Language Models
by: Lee, Tony, et al.
Published: (2025)
by: Lee, Tony, et al.
Published: (2025)
Chasing the Public Score: User Pressure and Evaluation Exploitation in Coding Agent Workflows
by: Chen, Hardy, et al.
Published: (2026)
by: Chen, Hardy, et al.
Published: (2026)
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness
by: Wang, Zeyu, et al.
Published: (2025)
by: Wang, Zeyu, et al.
Published: (2025)
SpatialThinker: Reinforcing 3D Reasoning in Multimodal LLMs via Spatial Rewards
by: Batra, Hunar, et al.
Published: (2025)
by: Batra, Hunar, et al.
Published: (2025)
MedVLThinker: Simple Baselines for Multimodal Medical Reasoning
by: Huang, Xiaoke, et al.
Published: (2025)
by: Huang, Xiaoke, et al.
Published: (2025)
VLAA-GUI: Knowing When to Stop, Recover, and Search, A Modular Framework for GUI Automation
by: Han, Qijun, et al.
Published: (2026)
by: Han, Qijun, et al.
Published: (2026)
OpenVision 3: A Family of Unified Visual Encoder for Both Understanding and Generation
by: Zhang, Letian, et al.
Published: (2026)
by: Zhang, Letian, et al.
Published: (2026)
When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought
by: Zhou, Yiyang, et al.
Published: (2025)
by: Zhou, Yiyang, et al.
Published: (2025)
On the Adversarial Robustness of Camera-based 3D Object Detection
by: Xie, Shaoyuan, et al.
Published: (2023)
by: Xie, Shaoyuan, et al.
Published: (2023)
Fantastic Bugs and Where to Find Them in AI Benchmarks
by: Truong, Sang, et al.
Published: (2025)
by: Truong, Sang, et al.
Published: (2025)
Medical Vision Generalist: Unifying Medical Imaging Tasks in Context
by: Ren, Sucheng, et al.
Published: (2024)
by: Ren, Sucheng, et al.
Published: (2024)
MetaClaw: Just Talk -- An Agent That Meta-Learns and Evolves in the Wild
by: Xia, Peng, et al.
Published: (2026)
by: Xia, Peng, et al.
Published: (2026)
Scaling White-Box Transformers for Vision
by: Yang, Jinrui, et al.
Published: (2024)
by: Yang, Jinrui, et al.
Published: (2024)
A Preliminary Study of o1 in Medicine: Are We Closer to an AI Doctor?
by: Xie, Yunfei, et al.
Published: (2024)
by: Xie, Yunfei, et al.
Published: (2024)
$\texttt{Complex-Edit}$: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark
by: Yang, Siwei, et al.
Published: (2025)
by: Yang, Siwei, et al.
Published: (2025)
STAR-1: Safer Alignment of Reasoning LLMs with 1K Data
by: Wang, Zijun, et al.
Published: (2025)
by: Wang, Zijun, et al.
Published: (2025)
Target-Oriented Pretraining Data Selection via Neuron-Activated Graph
by: Wang, Zijun, et al.
Published: (2026)
by: Wang, Zijun, et al.
Published: (2026)
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
by: Liu, Yanqing, et al.
Published: (2024)
by: Liu, Yanqing, et al.
Published: (2024)
LightFusion: A Light-weighted, Double Fusion Framework for Unified Multimodal Understanding and Generation
by: Wang, Zeyu, et al.
Published: (2025)
by: Wang, Zeyu, et al.
Published: (2025)
Language Models Can See Better: Visual Contrastive Decoding For LLM Multimodal Reasoning
by: Pang, Yuqi, et al.
Published: (2025)
by: Pang, Yuqi, et al.
Published: (2025)
Where Do Vision-Language Models Fail? World Scale Analysis for Image Geolocalization
by: Bharadwaj, Siddhant, et al.
Published: (2026)
by: Bharadwaj, Siddhant, et al.
Published: (2026)
Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More
by: Wang, Feng, et al.
Published: (2025)
by: Wang, Feng, et al.
Published: (2025)
Autoregressive Pretraining with Mamba in Vision
by: Ren, Sucheng, et al.
Published: (2024)
by: Ren, Sucheng, et al.
Published: (2024)
SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
by: Xia, Peng, et al.
Published: (2026)
by: Xia, Peng, et al.
Published: (2026)
Similar Items
-
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
by: Chen, Hardy, et al.
Published: (2025) -
ViLBench: A Suite for Vision-Language Process Reward Modeling
by: Tu, Haoqin, et al.
Published: (2025) -
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
by: Wu, Juncheng, et al.
Published: (2026) -
ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoning
by: Wu, Juncheng, et al.
Published: (2026) -
Kestrel: Grounding Self-Refinement for LVLM Hallucination Mitigation
by: Mao, Jiawei, et al.
Published: (2026)