Do Large Language Models Truly Understand Geometric Structures?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Xiaofeng, Wang, Yiming, Zhu, Wenhong, Wang, Rui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909503869419520
author Wang, Xiaofeng
Wang, Yiming
Zhu, Wenhong
Wang, Rui
author_facet Wang, Xiaofeng
Wang, Yiming
Zhu, Wenhong
Wang, Rui
contents Geometric ability is a significant challenge for large language models (LLMs) due to the need for advanced spatial comprehension and abstract thinking. Existing datasets primarily evaluate LLMs on their final answers, but they cannot truly measure their true understanding of geometric structures, as LLMs can arrive at correct answers by coincidence. To fill this gap, we introduce the GeomRel dataset, designed to evaluate LLMs' understanding of geometric structures by isolating the core step of geometric relationship identification in problem-solving. Using this benchmark, we conduct thorough evaluations of diverse LLMs and identify key limitations in understanding geometric structures. We further propose the Geometry Chain-of-Thought (GeoCoT) method, which enhances LLMs' ability to identify geometric relationships, resulting in significant performance improvements.
format Preprint
id arxiv_https___arxiv_org_abs_2501_13773
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Do Large Language Models Truly Understand Geometric Structures?
Wang, Xiaofeng
Wang, Yiming
Zhu, Wenhong
Wang, Rui
Computation and Language
Geometric ability is a significant challenge for large language models (LLMs) due to the need for advanced spatial comprehension and abstract thinking. Existing datasets primarily evaluate LLMs on their final answers, but they cannot truly measure their true understanding of geometric structures, as LLMs can arrive at correct answers by coincidence. To fill this gap, we introduce the GeomRel dataset, designed to evaluate LLMs' understanding of geometric structures by isolating the core step of geometric relationship identification in problem-solving. Using this benchmark, we conduct thorough evaluations of diverse LLMs and identify key limitations in understanding geometric structures. We further propose the Geometry Chain-of-Thought (GeoCoT) method, which enhances LLMs' ability to identify geometric relationships, resulting in significant performance improvements.
title Do Large Language Models Truly Understand Geometric Structures?
topic Computation and Language
url https://arxiv.org/abs/2501.13773