GraphOmni: A Comprehensive and Extensible Benchmark Framework for Large Language Models on Graph-theoretic Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Hao, Jian, Xiangru, Zhao, Xinjian, Pang, Wei, Zhang, Chao, Wang, Suyuchen, Zhang, Qixin, Dong, Zhengyuan, Monteiro, Joao, Liu, Bang, Sun, Qiuzhuang, Yu, Tianshu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918407857766400
author Xu, Hao
Jian, Xiangru
Zhao, Xinjian
Pang, Wei
Zhang, Chao
Wang, Suyuchen
Zhang, Qixin
Dong, Zhengyuan
Monteiro, Joao
Liu, Bang
Sun, Qiuzhuang
Yu, Tianshu
author_facet Xu, Hao
Jian, Xiangru
Zhao, Xinjian
Pang, Wei
Zhang, Chao
Wang, Suyuchen
Zhang, Qixin
Dong, Zhengyuan
Monteiro, Joao
Liu, Bang
Sun, Qiuzhuang
Yu, Tianshu
contents This paper introduces GraphOmni, a comprehensive benchmark designed to evaluate the reasoning capabilities of LLMs on graph-theoretic tasks articulated in natural language. GraphOmni encompasses diverse graph types, serialization formats, and prompting schemes, significantly exceeding prior efforts in both scope and depth. Through extensive systematic evaluation, we identify critical interactions among these dimensions, demonstrating their substantial impact on model performance. Our experiments reveal that state-of-the-art models like Claude-3.5 and o4-mini consistently outperform other models, yet even these leading models exhibit substantial room for improvement. Performance variability is evident depending on the specific combinations of factors we considered, underscoring the necessity of comprehensive evaluations across these interconnected dimensions. Additionally, we observe distinct impacts of serialization and prompting strategies between open-source and closed-source models, encouraging the development of tailored approaches. Motivated by the findings, we also propose a reinforcement learning-inspired framework that adaptively selects the optimal factors influencing LLM reasoning capabilities. This flexible and extendable benchmark not only deepens our understanding of LLM performance on structured tasks but also provides a robust foundation for advancing research in LLM-based graph reasoning. The code and datasets are available at https://github.com/GAI-Community/GraphOmni.
format Preprint
id arxiv_https___arxiv_org_abs_2504_12764
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GraphOmni: A Comprehensive and Extensible Benchmark Framework for Large Language Models on Graph-theoretic Tasks
Xu, Hao
Jian, Xiangru
Zhao, Xinjian
Pang, Wei
Zhang, Chao
Wang, Suyuchen
Zhang, Qixin
Dong, Zhengyuan
Monteiro, Joao
Liu, Bang
Sun, Qiuzhuang
Yu, Tianshu
Machine Learning
Discrete Mathematics
This paper introduces GraphOmni, a comprehensive benchmark designed to evaluate the reasoning capabilities of LLMs on graph-theoretic tasks articulated in natural language. GraphOmni encompasses diverse graph types, serialization formats, and prompting schemes, significantly exceeding prior efforts in both scope and depth. Through extensive systematic evaluation, we identify critical interactions among these dimensions, demonstrating their substantial impact on model performance. Our experiments reveal that state-of-the-art models like Claude-3.5 and o4-mini consistently outperform other models, yet even these leading models exhibit substantial room for improvement. Performance variability is evident depending on the specific combinations of factors we considered, underscoring the necessity of comprehensive evaluations across these interconnected dimensions. Additionally, we observe distinct impacts of serialization and prompting strategies between open-source and closed-source models, encouraging the development of tailored approaches. Motivated by the findings, we also propose a reinforcement learning-inspired framework that adaptively selects the optimal factors influencing LLM reasoning capabilities. This flexible and extendable benchmark not only deepens our understanding of LLM performance on structured tasks but also provides a robust foundation for advancing research in LLM-based graph reasoning. The code and datasets are available at https://github.com/GAI-Community/GraphOmni.
title GraphOmni: A Comprehensive and Extensible Benchmark Framework for Large Language Models on Graph-theoretic Tasks
topic Machine Learning
Discrete Mathematics
url https://arxiv.org/abs/2504.12764