Over-Reasoning and Redundant Calculation of Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chiang, Cheng-Han, Lee, Hung-yi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929282488467456
author Chiang, Cheng-Han
Lee, Hung-yi
author_facet Chiang, Cheng-Han
Lee, Hung-yi
contents Large language models (LLMs) can solve problems step-by-step. While this chain-of-thought (CoT) reasoning boosts LLMs' performance, it is unclear if LLMs \textit{know} when to use CoT and whether those CoT are always necessary to answer the question. This paper shows that LLMs tend to generate redundant calculations and reasoning on a manually constructed math QA dataset, GSM8K-Zero. GSM8K-Zero is constructed such that the questions can be answered without any calculations, but LLMs, including Llama-2 models and Claude-2, tend to generate lengthy and unnecessary calculations to answer the questions. We also conduct experiments to explain why LLMs generate redundant calculations and reasonings. GSM8K-Zero is publicly available at https://github.com/d223302/Over-Reasoning-of-LLMs and https://huggingface.co/datasets/dcml0714/GSM8K-Zero.
format Preprint
id arxiv_https___arxiv_org_abs_2401_11467
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Over-Reasoning and Redundant Calculation of Large Language Models
Chiang, Cheng-Han
Lee, Hung-yi
Computation and Language
Large language models (LLMs) can solve problems step-by-step. While this chain-of-thought (CoT) reasoning boosts LLMs' performance, it is unclear if LLMs \textit{know} when to use CoT and whether those CoT are always necessary to answer the question. This paper shows that LLMs tend to generate redundant calculations and reasoning on a manually constructed math QA dataset, GSM8K-Zero. GSM8K-Zero is constructed such that the questions can be answered without any calculations, but LLMs, including Llama-2 models and Claude-2, tend to generate lengthy and unnecessary calculations to answer the questions. We also conduct experiments to explain why LLMs generate redundant calculations and reasonings. GSM8K-Zero is publicly available at https://github.com/d223302/Over-Reasoning-of-LLMs and https://huggingface.co/datasets/dcml0714/GSM8K-Zero.
title Over-Reasoning and Redundant Calculation of Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2401.11467