GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces
Fuente:
arXiv
Saved in:
| Main Authors: | Geng, Xinyu, Xiao, Yanjing, Zhang, Yuyang, Wang, Hanwen, Liu, Xinyan, Min, Rui, Fang, Tianqing, Fung, Yi R. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization
by: Wang, Yikun, et al.
Published: (2025)
by: Wang, Yikun, et al.
Published: (2025)
GeoRC: A Benchmark for Geolocation Reasoning Chains
by: Talreja, Mohit, et al.
Published: (2026)
by: Talreja, Mohit, et al.
Published: (2026)
MedBrowseComp: Benchmarking Medical Deep Research and Computer Use
by: Chen, Shan, et al.
Published: (2025)
by: Chen, Shan, et al.
Published: (2025)
BrowseMaster: Towards Scalable Web Browsing via Tool-Augmented Programmatic Agent Pair
by: Pang, Xianghe, et al.
Published: (2025)
by: Pang, Xianghe, et al.
Published: (2025)
A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videos
by: Kurpath, Mohammed Irfan, et al.
Published: (2025)
by: Kurpath, Mohammed Irfan, et al.
Published: (2025)
ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning
by: Ding, Shengyuan, et al.
Published: (2025)
by: Ding, Shengyuan, et al.
Published: (2025)
MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents
by: Tao, Xijia, et al.
Published: (2025)
by: Tao, Xijia, et al.
Published: (2025)
VisBrowse-Bench: Benchmarking Visual-Native Search for Multimodal Browsing Agents
by: Zhang, Zhengbo, et al.
Published: (2026)
by: Zhang, Zhengbo, et al.
Published: (2026)
Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools
by: Wu, Junde, et al.
Published: (2025)
by: Wu, Junde, et al.
Published: (2025)
A Tool to Facilitate Web-Browsing
by: Kelly, Christopher, et al.
Published: (2024)
by: Kelly, Christopher, et al.
Published: (2024)
MCPVerse: An Expansive, Real-World Benchmark for Agentic Tool Use
by: Lei, Fei, et al.
Published: (2025)
by: Lei, Fei, et al.
Published: (2025)
BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
by: Wei, Jason, et al.
Published: (2025)
by: Wei, Jason, et al.
Published: (2025)
MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
by: Li, Shilong, et al.
Published: (2025)
by: Li, Shilong, et al.
Published: (2025)
Agentic Tool Use in Large Language Models
by: Hu, Jinchao, et al.
Published: (2026)
by: Hu, Jinchao, et al.
Published: (2026)
ToolRM: Towards Agentic Tool-Use Reward Modeling
by: Li, Renhao, et al.
Published: (2025)
by: Li, Renhao, et al.
Published: (2025)
AI Knowledge and Reasoning: Emulating Expert Creativity in Scientific Research
by: Mukherjee, Anirban, et al.
Published: (2024)
by: Mukherjee, Anirban, et al.
Published: (2024)
Browsing Lost Unformed Recollections: A Benchmark for Tip-of-the-Tongue Search and Reasoning
by: CH-Wang, Sky, et al.
Published: (2025)
by: CH-Wang, Sky, et al.
Published: (2025)
Tools as Continuous Flow for Evolving Agentic Reasoning
by: Huang, Tairan, et al.
Published: (2026)
by: Huang, Tairan, et al.
Published: (2026)
MedVR: Annotation-Free Medical Visual Reasoning via Agentic Reinforcement Learning
by: Jiang, Zheng, et al.
Published: (2026)
by: Jiang, Zheng, et al.
Published: (2026)
BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese
by: Zhou, Peilin, et al.
Published: (2025)
by: Zhou, Peilin, et al.
Published: (2025)
K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts
by: Lee, Nahyun, et al.
Published: (2026)
by: Lee, Nahyun, et al.
Published: (2026)
Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs
by: Lu, Meng, et al.
Published: (2025)
by: Lu, Meng, et al.
Published: (2025)
Video-Browser: Towards Agentic Open-web Video Browsing
by: Liang, Zhengyang, et al.
Published: (2025)
by: Liang, Zhengyang, et al.
Published: (2025)
Branch-and-Browse: Efficient and Controllable Web Exploration with Tree-Structured Reasoning and Action Memory
by: He, Shiqi, et al.
Published: (2025)
by: He, Shiqi, et al.
Published: (2025)
TRAJECT-Bench:A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use
by: He, Pengfei, et al.
Published: (2025)
by: He, Pengfei, et al.
Published: (2025)
CKBP v2: Better Annotation and Reasoning for Commonsense Knowledge Base Population
by: Fang, Tianqing, et al.
Published: (2023)
by: Fang, Tianqing, et al.
Published: (2023)
BrowseComp-$V^3$: A Visual, Vertical, and Verifiable Benchmark for Multimodal Browsing Agents
by: Zhang, Huanyao, et al.
Published: (2026)
by: Zhang, Huanyao, et al.
Published: (2026)
On Browsing: The Use of Search Theory in the Search for Information.
by: Morse, Philip M.
Published: (1970)
by: Morse, Philip M.
Published: (1970)
GeoMind: An Agentic Workflow for Lithology Classification with Reasoned Tool Invocation
by: Zhou, Yitong, et al.
Published: (2026)
by: Zhou, Yitong, et al.
Published: (2026)
ReasoningShield: Safety Detection over Reasoning Traces of Large Reasoning Models
by: Li, Changyi, et al.
Published: (2025)
by: Li, Changyi, et al.
Published: (2025)
VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
by: Jiang, Dongfu, et al.
Published: (2025)
by: Jiang, Dongfu, et al.
Published: (2025)
Heuristic Reasoning in AI: Instrumental Use and Mimetic Absorption
by: Mukherjee, Anirban, et al.
Published: (2024)
by: Mukherjee, Anirban, et al.
Published: (2024)
PaperArena: An Evaluation Benchmark for Tool-Augmented Agentic Reasoning on Scientific Literature
by: Wang, Daoyu, et al.
Published: (2025)
by: Wang, Daoyu, et al.
Published: (2025)
Vision-Language Reasoning for Geolocalization: A Reinforcement Learning Approach
by: Wu, Biao, et al.
Published: (2026)
by: Wu, Biao, et al.
Published: (2026)
On Stable Long-Form Generation: Benchmarking and Mitigating Length Volatility
by: He, Zhitao, et al.
Published: (2026)
by: He, Zhitao, et al.
Published: (2026)
Benchmarking LLM Tool-Use in the Wild
by: Yu, Peijie, et al.
Published: (2026)
by: Yu, Peijie, et al.
Published: (2026)
Interpretable Perception and Reasoning for Audiovisual Geolocation
by: Su, Yiyang, et al.
Published: (2026)
by: Su, Yiyang, et al.
Published: (2026)
Phantom Use: Quantifying In-Library Browsing of Circulating Materials
by: Wagner, Victoria H.
Published: (2007)
by: Wagner, Victoria H.
Published: (2007)
SCAN: Self-Denoising Monte Carlo Annotation for Robust Process Reward Learning
by: Ding, Yuyang, et al.
Published: (2025)
by: Ding, Yuyang, et al.
Published: (2025)
TRAIL: Trace Reasoning and Agentic Issue Localization
by: Deshpande, Darshan, et al.
Published: (2025)
by: Deshpande, Darshan, et al.
Published: (2025)
Similar Items
-
GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization
by: Wang, Yikun, et al.
Published: (2025) -
GeoRC: A Benchmark for Geolocation Reasoning Chains
by: Talreja, Mohit, et al.
Published: (2026) -
MedBrowseComp: Benchmarking Medical Deep Research and Computer Use
by: Chen, Shan, et al.
Published: (2025) -
BrowseMaster: Towards Scalable Web Browsing via Tool-Augmented Programmatic Agent Pair
by: Pang, Xianghe, et al.
Published: (2025) -
A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videos
by: Kurpath, Mohammed Irfan, et al.
Published: (2025)