Exploring the Compositional Deficiency of Large Language Models in Mathematical Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Jun, Tong, Jingqi, Mou, Yurong, Zhang, Ming, Zhang, Qi, Huang, Xuanjing
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914968804261888
author Zhao, Jun
Tong, Jingqi
Mou, Yurong
Zhang, Ming
Zhang, Qi
Huang, Xuanjing
author_facet Zhao, Jun
Tong, Jingqi
Mou, Yurong
Zhang, Ming
Zhang, Qi
Huang, Xuanjing
contents Human cognition exhibits systematic compositionality, the algebraic ability to generate infinite novel combinations from finite learned components, which is the key to understanding and reasoning about complex logic. In this work, we investigate the compositionality of large language models (LLMs) in mathematical reasoning. Specifically, we construct a new dataset \textsc{MathTrap} by introducing carefully designed logical traps into the problem descriptions of MATH and GSM8K. Since problems with logical flaws are quite rare in the real world, these represent "unseen" cases to LLMs. Solving these requires the models to systematically compose (1) the mathematical knowledge involved in the original problems with (2) knowledge related to the introduced traps. Our experiments show that while LLMs possess both components of requisite knowledge, they do not \textbf{spontaneously} combine them to handle these novel cases. We explore several methods to mitigate this deficiency, such as natural language prompts, few-shot demonstrations, and fine-tuning. Additionally, we test the recently released OpenAI o1 model and find that human-like `slow thinking' helps improve the compositionality of LLMs. Overall, systematic compositionality remains an open challenge for large language models.
format Preprint
id arxiv_https___arxiv_org_abs_2405_06680
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exploring the Compositional Deficiency of Large Language Models in Mathematical Reasoning
Zhao, Jun
Tong, Jingqi
Mou, Yurong
Zhang, Ming
Zhang, Qi
Huang, Xuanjing
Computation and Language
Artificial Intelligence
Human cognition exhibits systematic compositionality, the algebraic ability to generate infinite novel combinations from finite learned components, which is the key to understanding and reasoning about complex logic. In this work, we investigate the compositionality of large language models (LLMs) in mathematical reasoning. Specifically, we construct a new dataset \textsc{MathTrap} by introducing carefully designed logical traps into the problem descriptions of MATH and GSM8K. Since problems with logical flaws are quite rare in the real world, these represent "unseen" cases to LLMs. Solving these requires the models to systematically compose (1) the mathematical knowledge involved in the original problems with (2) knowledge related to the introduced traps. Our experiments show that while LLMs possess both components of requisite knowledge, they do not \textbf{spontaneously} combine them to handle these novel cases. We explore several methods to mitigate this deficiency, such as natural language prompts, few-shot demonstrations, and fine-tuning. Additionally, we test the recently released OpenAI o1 model and find that human-like `slow thinking' helps improve the compositionality of LLMs. Overall, systematic compositionality remains an open challenge for large language models.
title Exploring the Compositional Deficiency of Large Language Models in Mathematical Reasoning
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2405.06680