EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Wanghan, Zhao, Xiangyu, Zhou, Yuhao, Yue, Xiaoyu, Fei, Ben, Ling, Fenghua, Zhang, Wenlong, Bai, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MSEarth: A Multimodal Benchmark for Earth Science Phenomenon Discovery with MLLMs
by: Zhao, Xiangyu, et al.
Published: (2025)
by: Zhao, Xiangyu, et al.
Published: (2025)
Earth Science Foundation Models: From Perception to Reasoning and Discovery
by: Zhao, Xiangyu, et al.
Published: (2026)
by: Zhao, Xiangyu, et al.
Published: (2026)
ReCrit: Transition-Aware Reinforcement Learning for Scientific Critic Reasoning
by: Xu, Wanghan, et al.
Published: (2026)
by: Xu, Wanghan, et al.
Published: (2026)
Earth-Agent: Unlocking the Full Landscape of Earth Observation with Agents
by: Feng, Peilin, et al.
Published: (2025)
by: Feng, Peilin, et al.
Published: (2025)
OpenEarth-Agent: From Tool Calling to Tool Creation for Open-Environment Earth Observation
by: Zhao, Sijie, et al.
Published: (2026)
by: Zhao, Sijie, et al.
Published: (2026)
Generalizing Weather Forecast to Fine-grained Temporal Scales via Physics-AI Hybrid Modeling
by: Xu, Wanghan, et al.
Published: (2024)
by: Xu, Wanghan, et al.
Published: (2024)
DAWP: A framework for global observation forecasting via Data Assimilation and Weather Prediction in satellite observation space
by: Gong, Junchao, et al.
Published: (2025)
by: Gong, Junchao, et al.
Published: (2025)
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities
by: Li, Haoming, et al.
Published: (2025)
by: Li, Haoming, et al.
Published: (2025)
FOFO: A Benchmark to Evaluate LLMs' Format-Following Capability
by: Xia, Congying, et al.
Published: (2024)
by: Xia, Congying, et al.
Published: (2024)
EarthSpatialBench: Benchmarking Spatial Reasoning Capabilities of Multimodal LLMs on Earth Imagery
by: Xu, Zelin, et al.
Published: (2026)
by: Xu, Zelin, et al.
Published: (2026)
SciEvalKit: An Open-source Evaluation Toolkit for Scientific General Intelligence
by: Wang, Yiheng, et al.
Published: (2025)
by: Wang, Yiheng, et al.
Published: (2025)
LLMsPark: A Benchmark for Evaluating Large Language Models in Strategic Gaming Contexts
by: Chen, Junhao, et al.
Published: (2025)
by: Chen, Junhao, et al.
Published: (2025)
Evaluating LLMs' Divergent Thinking Capabilities for Scientific Idea Generation with Minimal Context
by: Ruan, Kai, et al.
Published: (2024)
by: Ruan, Kai, et al.
Published: (2024)
Manalyzer: End-to-end Automated Meta-analysis with Multi-agent System
by: Xu, Wanghan, et al.
Published: (2025)
by: Xu, Wanghan, et al.
Published: (2025)
Align-DA: Align Score-based Atmospheric Data Assimilation with Multiple Preferences
by: Sun, Jing-An, et al.
Published: (2025)
by: Sun, Jing-An, et al.
Published: (2025)
LO-SDA: Latent Optimization for Score-based Atmospheric Data Assimilation
by: Sun, Jing-An, et al.
Published: (2025)
by: Sun, Jing-An, et al.
Published: (2025)
Eigen-1: Adaptive Multi-Agent Refinement with Monitor-Based RAG for Scientific Reasoning
by: Tang, Xiangru, et al.
Published: (2025)
by: Tang, Xiangru, et al.
Published: (2025)
OmniEarth-Bench: Towards Holistic Evaluation of Earth's Six Spheres and Cross-Spheres Interactions with Multimodal Observational Earth Data
by: Wang, Fengxiang, et al.
Published: (2025)
by: Wang, Fengxiang, et al.
Published: (2025)
SciReasoner: Laying the Scientific Reasoning Ground Across Disciplines
by: Wang, Yizhou, et al.
Published: (2025)
by: Wang, Yizhou, et al.
Published: (2025)
Evaluating LLMs' Multilingual Capabilities for Bengali: Benchmark Creation and Performance Analysis
by: Bhowmik, Shimanto, et al.
Published: (2025)
by: Bhowmik, Shimanto, et al.
Published: (2025)
Auto-SLURP: A Benchmark Dataset for Evaluating Multi-Agent Frameworks in Smart Personal Assistant
by: Shen, Lei, et al.
Published: (2025)
by: Shen, Lei, et al.
Published: (2025)
Meeseeks: A Feedback-Driven, Iterative Self-Correction Benchmark evaluating LLMs' Instruction Following Capability
by: wang, Jiaming, et al.
Published: (2025)
by: wang, Jiaming, et al.
Published: (2025)
OpenEval: Benchmarking Chinese LLMs across Capability, Alignment and Safety
by: Liu, Chuang, et al.
Published: (2024)
by: Liu, Chuang, et al.
Published: (2024)
Towards Automatic Evaluation for LLMs' Clinical Capabilities: Metric, Data, and Algorithm
by: Liu, Lei, et al.
Published: (2024)
by: Liu, Lei, et al.
Published: (2024)
CCiV: A Benchmark for Structure, Rhythm and Quality in LLM-Generated Chinese \textit{Ci} Poetry
by: Zhao, Shangqing, et al.
Published: (2026)
by: Zhao, Shangqing, et al.
Published: (2026)
Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs
by: Son, Guijin, et al.
Published: (2026)
by: Son, Guijin, et al.
Published: (2026)
LabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs
by: Zhou, Yujun, et al.
Published: (2024)
by: Zhou, Yujun, et al.
Published: (2024)
P-MMEval: A Parallel Multilingual Multitask Benchmark for Consistent Evaluation of LLMs
by: Zhang, Yidan, et al.
Published: (2024)
by: Zhang, Yidan, et al.
Published: (2024)
Can LLMs Act as Historians? Evaluating Historical Research Capabilities of LLMs via the Chinese Imperial Examination
by: Gao, Lirong, et al.
Published: (2026)
by: Gao, Lirong, et al.
Published: (2026)
Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows
by: Xu, Wanghan, et al.
Published: (2025)
by: Xu, Wanghan, et al.
Published: (2025)
Korean Canonical Legal Benchmark: Toward Knowledge-Independent Evaluation of LLMs' Legal Reasoning Capabilities
by: Oh, Hongseok, et al.
Published: (2025)
by: Oh, Hongseok, et al.
Published: (2025)
IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages
by: Singh, Harman, et al.
Published: (2024)
by: Singh, Harman, et al.
Published: (2024)
DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding
by: Zhu, Hengchuan, et al.
Published: (2025)
by: Zhu, Hengchuan, et al.
Published: (2025)
Be Cautious When Merging Unfamiliar LLMs: A Phishing Model Capable of Stealing Privacy
by: Guo, Zhenyuan, et al.
Published: (2025)
by: Guo, Zhenyuan, et al.
Published: (2025)
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
by: Jiang, Zhuohang, et al.
Published: (2025)
by: Jiang, Zhuohang, et al.
Published: (2025)
The Earth is Flat because...: Investigating LLMs' Belief towards Misinformation via Persuasive Conversation
by: Xu, Rongwu, et al.
Published: (2023)
by: Xu, Rongwu, et al.
Published: (2023)
Unlocking Reasoning Capabilities in LLMs via Reinforcement Learning Exploration
by: Deng, Wenhao, et al.
Published: (2025)
by: Deng, Wenhao, et al.
Published: (2025)
CHBench: A Cognitive Hierarchy Benchmark for Evaluating Strategic Reasoning Capability of LLMs
by: Liu, Hongtao, et al.
Published: (2025)
by: Liu, Hongtao, et al.
Published: (2025)
Benchmark Illusion: Disagreement among LLMs and Its Scientific Consequences
by: Yang, Eddie, et al.
Published: (2026)
by: Yang, Eddie, et al.
Published: (2026)
CharacterBox: Evaluating the Role-Playing Capabilities of LLMs in Text-Based Virtual Worlds
by: Wang, Lei, et al.
Published: (2024)
by: Wang, Lei, et al.
Published: (2024)
Similar Items
-
MSEarth: A Multimodal Benchmark for Earth Science Phenomenon Discovery with MLLMs
by: Zhao, Xiangyu, et al.
Published: (2025) -
Earth Science Foundation Models: From Perception to Reasoning and Discovery
by: Zhao, Xiangyu, et al.
Published: (2026) -
ReCrit: Transition-Aware Reinforcement Learning for Scientific Critic Reasoning
by: Xu, Wanghan, et al.
Published: (2026) -
Earth-Agent: Unlocking the Full Landscape of Earth Observation with Agents
by: Feng, Peilin, et al.
Published: (2025) -
OpenEarth-Agent: From Tool Calling to Tool Creation for Open-Environment Earth Observation
by: Zhao, Sijie, et al.
Published: (2026)