Saved in:
| Main Authors: | Xu, Ruiling, Zhang, Yifan, Wang, Qingyun, Edwards, Carl, Ji, Heng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2510.07731 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
L+M-24: Building a Dataset for Language + Molecules @ ACL 2024
by: Edwards, Carl, et al.
Published: (2024)
by: Edwards, Carl, et al.
Published: (2024)
CANDY: Benchmarking LLMs' Limitations and Assistive Potential in Chinese Misinformation Fact-Checking
by: Guo, Ruiling, et al.
Published: (2025)
by: Guo, Ruiling, et al.
Published: (2025)
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
by: Jiang, Zhuohang, et al.
Published: (2025)
by: Jiang, Zhuohang, et al.
Published: (2025)
TopoBench: Benchmarking LLMs on Hard Topological Reasoning
by: Maniparambil, Mayug, et al.
Published: (2026)
by: Maniparambil, Mayug, et al.
Published: (2026)
MolCap-Arena: A Comprehensive Captioning Benchmark on Language-Enhanced Molecular Property Prediction
by: Edwards, Carl, et al.
Published: (2024)
by: Edwards, Carl, et al.
Published: (2024)
LocalBench: Benchmarking LLMs on County-Level Local Knowledge and Reasoning
by: Gao, Zihan, et al.
Published: (2025)
by: Gao, Zihan, et al.
Published: (2025)
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
by: Kabir, Mohsinul, et al.
Published: (2026)
by: Kabir, Mohsinul, et al.
Published: (2026)
CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
by: Lin, Zicheng, et al.
Published: (2024)
by: Lin, Zicheng, et al.
Published: (2024)
MoralBench: Moral Evaluation of LLMs
by: Ji, Jianchao, et al.
Published: (2024)
by: Ji, Jianchao, et al.
Published: (2024)
SciRerankBench: Benchmarking Rerankers Towards Scientific Retrieval-Augmented Generated LLMs
by: Chen, Haotian, et al.
Published: (2025)
by: Chen, Haotian, et al.
Published: (2025)
SciMON: Scientific Inspiration Machines Optimized for Novelty
by: Wang, Qingyun, et al.
Published: (2023)
by: Wang, Qingyun, et al.
Published: (2023)
Stage-wise Fine-tuning for Graph-to-Text Generation
by: Wang, Qingyun, et al.
Published: (2021)
by: Wang, Qingyun, et al.
Published: (2021)
HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
by: Chen, Junying, et al.
Published: (2024)
by: Chen, Junying, et al.
Published: (2024)
AQA-Bench: An Interactive Benchmark for Evaluating LLMs' Sequential Reasoning Ability
by: Yang, Siwei, et al.
Published: (2024)
by: Yang, Siwei, et al.
Published: (2024)
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs
by: Kil, Jihyung, et al.
Published: (2024)
by: Kil, Jihyung, et al.
Published: (2024)
Explore the Reasoning Capability of LLMs in the Chess Testbed
by: Wang, Shu, et al.
Published: (2024)
by: Wang, Shu, et al.
Published: (2024)
MolQuest: A Benchmark for Agentic Evaluation of Abductive Reasoning in Chemical Structure Elucidation
by: Han, Taolin, et al.
Published: (2026)
by: Han, Taolin, et al.
Published: (2026)
DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding
by: Zhu, Hengchuan, et al.
Published: (2025)
by: Zhu, Hengchuan, et al.
Published: (2025)
seqBench: A Tunable Benchmark to Quantify Sequential Reasoning Limits of LLMs
by: Ramezanali, Mohammad, et al.
Published: (2025)
by: Ramezanali, Mohammad, et al.
Published: (2025)
Call Me When Necessary: LLMs can Efficiently and Faithfully Reason over Structured Environments
by: Cheng, Sitao, et al.
Published: (2024)
by: Cheng, Sitao, et al.
Published: (2024)
ArgBench: Benchmarking LLMs on Computational Argumentation Tasks
by: Ajjour, Yamen, et al.
Published: (2026)
by: Ajjour, Yamen, et al.
Published: (2026)
Mechanics of Learned Reasoning 1: TempoBench, A Benchmark for Interpretable Deconstruction of Reasoning System Performance
by: Holzer, Nikolaus, et al.
Published: (2025)
by: Holzer, Nikolaus, et al.
Published: (2025)
Scaling Laws for Predicting Downstream Performance in LLMs
by: Chen, Yangyi, et al.
Published: (2024)
by: Chen, Yangyi, et al.
Published: (2024)
FaithBench: A Diverse Hallucination Benchmark for Summarization by Modern LLMs
by: Bao, Forrest Sheng, et al.
Published: (2024)
by: Bao, Forrest Sheng, et al.
Published: (2024)
ReviewRobot: Explainable Paper Review Generation based on Knowledge Synthesis
by: Wang, Qingyun, et al.
Published: (2020)
by: Wang, Qingyun, et al.
Published: (2020)
Lie to Me: Knowledge Graphs for Robust Hallucination Self-Detection in LLMs
by: Kale, Sahil, et al.
Published: (2025)
by: Kale, Sahil, et al.
Published: (2025)
PsychiatryBench: A Multi-Task Benchmark for LLMs in Psychiatry
by: Fouda, Aya E., et al.
Published: (2025)
by: Fouda, Aya E., et al.
Published: (2025)
UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models
by: Xu, Xin, et al.
Published: (2025)
by: Xu, Xin, et al.
Published: (2025)
Automating Intervention Discovery from Scientific Literature: A Progressive Ontology Prompting and Dual-LLM Framework
by: Hu, Yuting, et al.
Published: (2024)
by: Hu, Yuting, et al.
Published: (2024)
Computation Mechanism Behind LLM Position Generalization
by: Han, Chi, et al.
Published: (2025)
by: Han, Chi, et al.
Published: (2025)
TopBench: A Benchmark for Implicit Prediction and Reasoning over Tabular Question Answering
by: Ji, An-Yang, et al.
Published: (2026)
by: Ji, An-Yang, et al.
Published: (2026)
WritingBench: A Comprehensive Benchmark for Generative Writing
by: Wu, Yuning, et al.
Published: (2025)
by: Wu, Yuning, et al.
Published: (2025)
T2S-Bench & Structure-of-Thought: Benchmarking and Prompting Comprehensive Text-to-Structure Reasoning
by: Wang, Qinsi, et al.
Published: (2026)
by: Wang, Qinsi, et al.
Published: (2026)
LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
by: Bai, Yushi, et al.
Published: (2024)
by: Bai, Yushi, et al.
Published: (2024)
AgentBench: Evaluating LLMs as Agents
by: Liu, Xiao, et al.
Published: (2023)
by: Liu, Xiao, et al.
Published: (2023)
PLawBench: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice
by: Shi, Yuzhen, et al.
Published: (2026)
by: Shi, Yuzhen, et al.
Published: (2026)
MemoryRewardBench: Benchmarking Reward Models for Long-Term Memory Management in Large Language Models
by: Tang, Zecheng, et al.
Published: (2026)
by: Tang, Zecheng, et al.
Published: (2026)
CricBench: A Multilingual Benchmark for Evaluating LLMs in Cricket Analytics
by: Agarwal, Parth, et al.
Published: (2025)
by: Agarwal, Parth, et al.
Published: (2025)
BizFinBench: A Business-Driven Real-World Financial Benchmark for Evaluating LLMs
by: Lu, Guilong, et al.
Published: (2025)
by: Lu, Guilong, et al.
Published: (2025)
SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia
by: Liu, Chaoqun, et al.
Published: (2025)
by: Liu, Chaoqun, et al.
Published: (2025)
Similar Items
-
L+M-24: Building a Dataset for Language + Molecules @ ACL 2024
by: Edwards, Carl, et al.
Published: (2024) -
CANDY: Benchmarking LLMs' Limitations and Assistive Potential in Chinese Misinformation Fact-Checking
by: Guo, Ruiling, et al.
Published: (2025) -
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
by: Jiang, Zhuohang, et al.
Published: (2025) -
TopoBench: Benchmarking LLMs on Hard Topological Reasoning
by: Maniparambil, Mayug, et al.
Published: (2026) -
MolCap-Arena: A Comprehensive Captioning Benchmark on Language-Enhanced Molecular Property Prediction
by: Edwards, Carl, et al.
Published: (2024)