Can Large Language Models Generate Geospatial Code?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hou, Shuyang, Shen, Zhangxiao, Liang, Jianyuan, Zhao, Anqi, Gui, Zhipeng, Li, Rui, Wu, Huayi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910655793070080
author Hou, Shuyang
Shen, Zhangxiao
Liang, Jianyuan
Zhao, Anqi
Gui, Zhipeng
Li, Rui
Wu, Huayi
author_facet Hou, Shuyang
Shen, Zhangxiao
Liang, Jianyuan
Zhao, Anqi
Gui, Zhipeng
Li, Rui
Wu, Huayi
contents With the growing demand for spatiotemporal data processing and geospatial modeling, automating geospatial code generation has become essential for productivity. Large language models (LLMs) show promise in code generation but face challenges like domain-specific knowledge gaps and "coding hallucinations." This paper introduces GeoCode-Eval (GCE), a framework for assessing LLMs' ability to generate geospatial code across three dimensions: "Cognition and Memory," "Comprehension and Interpretation," and "Innovation and Creation," distributed across eight capability levels. We developed a benchmark dataset, GeoCode-Bench, consisting of 5,000 multiple-choice, 1,500 fill-in-the-blank, 1,500 true/false questions, and 1,000 subjective tasks covering code summarization, generation, completion, and correction. Using GeoCode-Bench, we evaluated three commercial closed-source LLMs, four open-source general-purpose LLMs, and 14 specialized code generation models. We also conducted experiments on few-shot and zero-shot learning, Chain of Thought reasoning, and multi-round majority voting to measure their impact on geospatial code generation. Additionally, we fine-tuned the Code LLaMA-7B model using Google Earth Engine-related JavaScript, creating GEECode-GPT, and evaluated it on subjective tasks. Results show that constructing pre-training and instruction datasets significantly improves code generation, offering insights for optimizing LLMs in specific domains.
format Preprint
id arxiv_https___arxiv_org_abs_2410_09738
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Can Large Language Models Generate Geospatial Code?
Hou, Shuyang
Shen, Zhangxiao
Liang, Jianyuan
Zhao, Anqi
Gui, Zhipeng
Li, Rui
Wu, Huayi
Software Engineering
With the growing demand for spatiotemporal data processing and geospatial modeling, automating geospatial code generation has become essential for productivity. Large language models (LLMs) show promise in code generation but face challenges like domain-specific knowledge gaps and "coding hallucinations." This paper introduces GeoCode-Eval (GCE), a framework for assessing LLMs' ability to generate geospatial code across three dimensions: "Cognition and Memory," "Comprehension and Interpretation," and "Innovation and Creation," distributed across eight capability levels. We developed a benchmark dataset, GeoCode-Bench, consisting of 5,000 multiple-choice, 1,500 fill-in-the-blank, 1,500 true/false questions, and 1,000 subjective tasks covering code summarization, generation, completion, and correction. Using GeoCode-Bench, we evaluated three commercial closed-source LLMs, four open-source general-purpose LLMs, and 14 specialized code generation models. We also conducted experiments on few-shot and zero-shot learning, Chain of Thought reasoning, and multi-round majority voting to measure their impact on geospatial code generation. Additionally, we fine-tuned the Code LLaMA-7B model using Google Earth Engine-related JavaScript, creating GEECode-GPT, and evaluated it on subjective tasks. Results show that constructing pre-training and instruction datasets significantly improves code generation, offering insights for optimizing LLMs in specific domains.
title Can Large Language Models Generate Geospatial Code?
topic Software Engineering
url https://arxiv.org/abs/2410.09738