GeoBenchX: Benchmarking LLMs in Agent Solving Multistep Geospatial Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Krechetova, Varvara, Kochedykov, Denis |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ArgBench: Benchmarking LLMs on Computational Argumentation Tasks
von: Ajjour, Yamen, et al.
Veröffentlicht: (2026)
von: Ajjour, Yamen, et al.
Veröffentlicht: (2026)
PsychiatryBench: A Multi-Task Benchmark for LLMs in Psychiatry
von: Fouda, Aya E., et al.
Veröffentlicht: (2025)
von: Fouda, Aya E., et al.
Veröffentlicht: (2025)
Ko-MuSR: A Multistep Soft Reasoning Benchmark for LLMs Capable of Understanding Korean
von: Park, Chanwoo, et al.
Veröffentlicht: (2025)
von: Park, Chanwoo, et al.
Veröffentlicht: (2025)
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2024)
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2024)
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
von: Kabir, Mohsinul, et al.
Veröffentlicht: (2026)
von: Kabir, Mohsinul, et al.
Veröffentlicht: (2026)
VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
von: He, Wei, et al.
Veröffentlicht: (2025)
von: He, Wei, et al.
Veröffentlicht: (2025)
TaskBench: Benchmarking Large Language Models for Task Automation
von: Shen, Yongliang, et al.
Veröffentlicht: (2023)
von: Shen, Yongliang, et al.
Veröffentlicht: (2023)
GeoResponder: Towards Building Geospatial LLMs for Time-Critical Disaster Response
von: Zguir, Ahmed El Fekih, et al.
Veröffentlicht: (2025)
von: Zguir, Ahmed El Fekih, et al.
Veröffentlicht: (2025)
GeoEval: Benchmark for Evaluating LLMs and Multi-Modal Models on Geometry Problem-Solving
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2024)
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2024)
ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon Tasks
von: Song, Yuanyi, et al.
Veröffentlicht: (2025)
von: Song, Yuanyi, et al.
Veröffentlicht: (2025)
FB-Bench: A Fine-Grained Multi-Task Benchmark for Evaluating LLMs' Responsiveness to Human Feedback
von: Li, Youquan, et al.
Veröffentlicht: (2024)
von: Li, Youquan, et al.
Veröffentlicht: (2024)
TopoBench: Benchmarking LLMs on Hard Topological Reasoning
von: Maniparambil, Mayug, et al.
Veröffentlicht: (2026)
von: Maniparambil, Mayug, et al.
Veröffentlicht: (2026)
AgentBench: Evaluating LLMs as Agents
von: Liu, Xiao, et al.
Veröffentlicht: (2023)
von: Liu, Xiao, et al.
Veröffentlicht: (2023)
Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks
von: Pires, Ramon, et al.
Veröffentlicht: (2026)
von: Pires, Ramon, et al.
Veröffentlicht: (2026)
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
Mem2ActBench: A Benchmark for Evaluating Long-Term Memory Utilization in Task-Oriented Autonomous Agents
von: Shen, Yiting, et al.
Veröffentlicht: (2026)
von: Shen, Yiting, et al.
Veröffentlicht: (2026)
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
von: Long, Xiang, et al.
Veröffentlicht: (2026)
von: Long, Xiang, et al.
Veröffentlicht: (2026)
ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition
von: Liu, Yujie, et al.
Veröffentlicht: (2025)
von: Liu, Yujie, et al.
Veröffentlicht: (2025)
DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding
von: Zhu, Hengchuan, et al.
Veröffentlicht: (2025)
von: Zhu, Hengchuan, et al.
Veröffentlicht: (2025)
CricBench: A Multilingual Benchmark for Evaluating LLMs in Cricket Analytics
von: Agarwal, Parth, et al.
Veröffentlicht: (2025)
von: Agarwal, Parth, et al.
Veröffentlicht: (2025)
FaithBench: A Diverse Hallucination Benchmark for Summarization by Modern LLMs
von: Bao, Forrest Sheng, et al.
Veröffentlicht: (2024)
von: Bao, Forrest Sheng, et al.
Veröffentlicht: (2024)
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
von: Deng, Shihan, et al.
Veröffentlicht: (2024)
von: Deng, Shihan, et al.
Veröffentlicht: (2024)
ClawBench: Can AI Agents Complete Everyday Online Tasks?
von: Zhang, Yuxuan, et al.
Veröffentlicht: (2026)
von: Zhang, Yuxuan, et al.
Veröffentlicht: (2026)
MedAgentsBench: Benchmarking Thinking Models and Agent Frameworks for Complex Medical Reasoning
von: Tang, Xiangru, et al.
Veröffentlicht: (2025)
von: Tang, Xiangru, et al.
Veröffentlicht: (2025)
FHIR-AgentBench: Benchmarking LLM Agents for Realistic Interoperable EHR Question Answering
von: Lee, Gyubok, et al.
Veröffentlicht: (2025)
von: Lee, Gyubok, et al.
Veröffentlicht: (2025)
GraphicBench: A Planning Benchmark for Graphic Design with Language Agents
von: Ki, Dayeon, et al.
Veröffentlicht: (2025)
von: Ki, Dayeon, et al.
Veröffentlicht: (2025)
MASLegalBench: Benchmarking Multi-Agent Systems in Deductive Legal Reasoning
von: Jing, Huihao, et al.
Veröffentlicht: (2025)
von: Jing, Huihao, et al.
Veröffentlicht: (2025)
ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences
von: Nguyen, Bang, et al.
Veröffentlicht: (2026)
von: Nguyen, Bang, et al.
Veröffentlicht: (2026)
KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs
von: Kim, Haechan, et al.
Veröffentlicht: (2026)
von: Kim, Haechan, et al.
Veröffentlicht: (2026)
ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks
von: Li, Minghao, et al.
Veröffentlicht: (2025)
von: Li, Minghao, et al.
Veröffentlicht: (2025)
SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
von: Orlanski, Gabriel, et al.
Veröffentlicht: (2026)
von: Orlanski, Gabriel, et al.
Veröffentlicht: (2026)
LocalBench: Benchmarking LLMs on County-Level Local Knowledge and Reasoning
von: Gao, Zihan, et al.
Veröffentlicht: (2025)
von: Gao, Zihan, et al.
Veröffentlicht: (2025)
CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
von: Lin, Zicheng, et al.
Veröffentlicht: (2024)
von: Lin, Zicheng, et al.
Veröffentlicht: (2024)
QuantumBench: A Benchmark for Quantum Problem Solving
von: Minami, Shunya, et al.
Veröffentlicht: (2025)
von: Minami, Shunya, et al.
Veröffentlicht: (2025)
BenchAgents: Multi-Agent Systems for Structured Benchmark Creation
von: Butt, Natasha, et al.
Veröffentlicht: (2024)
von: Butt, Natasha, et al.
Veröffentlicht: (2024)
ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code
von: Hua, Tianyu, et al.
Veröffentlicht: (2025)
von: Hua, Tianyu, et al.
Veröffentlicht: (2025)
SciRerankBench: Benchmarking Rerankers Towards Scientific Retrieval-Augmented Generated LLMs
von: Chen, Haotian, et al.
Veröffentlicht: (2025)
von: Chen, Haotian, et al.
Veröffentlicht: (2025)
SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia
von: Liu, Chaoqun, et al.
Veröffentlicht: (2025)
von: Liu, Chaoqun, et al.
Veröffentlicht: (2025)
SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs
von: Liu, Zhiqiang, et al.
Veröffentlicht: (2025)
von: Liu, Zhiqiang, et al.
Veröffentlicht: (2025)
oMeBench: Towards Robust Benchmarking of LLMs in Organic Mechanism Elucidation and Reasoning
von: Xu, Ruiling, et al.
Veröffentlicht: (2025)
von: Xu, Ruiling, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ArgBench: Benchmarking LLMs on Computational Argumentation Tasks
von: Ajjour, Yamen, et al.
Veröffentlicht: (2026) -
PsychiatryBench: A Multi-Task Benchmark for LLMs in Psychiatry
von: Fouda, Aya E., et al.
Veröffentlicht: (2025) -
Ko-MuSR: A Multistep Soft Reasoning Benchmark for LLMs Capable of Understanding Korean
von: Park, Chanwoo, et al.
Veröffentlicht: (2025) -
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2024) -
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
von: Kabir, Mohsinul, et al.
Veröffentlicht: (2026)