GCoder: Improving Large Language Model for Generalized Graph Problem Solving

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Qifan, Hong, Xiaobin, Tang, Jianheng, Chen, Nuo, Li, Yuhan, Li, Wenzhong, Tang, Jing, Li, Jia
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929557965111296
author Zhang, Qifan
Hong, Xiaobin
Tang, Jianheng
Chen, Nuo
Li, Yuhan
Li, Wenzhong
Tang, Jing
Li, Jia
author_facet Zhang, Qifan
Hong, Xiaobin
Tang, Jianheng
Chen, Nuo
Li, Yuhan
Li, Wenzhong
Tang, Jing
Li, Jia
contents Large Language Models (LLMs) have demonstrated strong reasoning abilities, making them suitable for complex tasks such as graph computation. Traditional reasoning steps paradigm for graph problems is hindered by unverifiable steps, limited long-term reasoning, and poor generalization to graph variations. To overcome these limitations, we introduce GCoder, a code-based LLM designed to enhance problem-solving in generalized graph computation problems. Our method involves constructing an extensive training dataset, GraphWild, featuring diverse graph formats and algorithms. We employ a multi-stage training process, including Supervised Fine-Tuning (SFT) and Reinforcement Learning from Compiler Feedback (RLCF), to refine model capabilities. For unseen tasks, a hybrid retrieval technique is used to augment performance. Experiments demonstrate that GCoder outperforms GPT-4o, with an average accuracy improvement of 16.42% across various graph computational problems. Furthermore, GCoder efficiently manages large-scale graphs with millions of nodes and diverse input formats, overcoming the limitations of previous models focused on the reasoning steps paradigm. This advancement paves the way for more intuitive and effective graph problem-solving using LLMs. Code and data are available at here: https://github.com/Bklight999/WWW25-GCoder/tree/master.
format Preprint
id arxiv_https___arxiv_org_abs_2410_19084
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GCoder: Improving Large Language Model for Generalized Graph Problem Solving
Zhang, Qifan
Hong, Xiaobin
Tang, Jianheng
Chen, Nuo
Li, Yuhan
Li, Wenzhong
Tang, Jing
Li, Jia
Computation and Language
Large Language Models (LLMs) have demonstrated strong reasoning abilities, making them suitable for complex tasks such as graph computation. Traditional reasoning steps paradigm for graph problems is hindered by unverifiable steps, limited long-term reasoning, and poor generalization to graph variations. To overcome these limitations, we introduce GCoder, a code-based LLM designed to enhance problem-solving in generalized graph computation problems. Our method involves constructing an extensive training dataset, GraphWild, featuring diverse graph formats and algorithms. We employ a multi-stage training process, including Supervised Fine-Tuning (SFT) and Reinforcement Learning from Compiler Feedback (RLCF), to refine model capabilities. For unseen tasks, a hybrid retrieval technique is used to augment performance. Experiments demonstrate that GCoder outperforms GPT-4o, with an average accuracy improvement of 16.42% across various graph computational problems. Furthermore, GCoder efficiently manages large-scale graphs with millions of nodes and diverse input formats, overcoming the limitations of previous models focused on the reasoning steps paradigm. This advancement paves the way for more intuitive and effective graph problem-solving using LLMs. Code and data are available at here: https://github.com/Bklight999/WWW25-GCoder/tree/master.
title GCoder: Improving Large Language Model for Generalized Graph Problem Solving
topic Computation and Language
url https://arxiv.org/abs/2410.19084