Understanding the Language Model to Solve the Symbolic Multi-Step Reasoning Problem from the Perspective of Buffer Mechanism

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zhiwei, Wang, Yunji, Zhang, Zhongwang, Zhou, Zhangchen, Jin, Hui, Hu, Tianyang, Sun, Jiacheng, Li, Zhenguo, Zhang, Yaoyu, Xu, Zhi-Qin John
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909777148248064
author Wang, Zhiwei
Wang, Yunji
Zhang, Zhongwang
Zhou, Zhangchen
Jin, Hui
Hu, Tianyang
Sun, Jiacheng
Li, Zhenguo
Zhang, Yaoyu
Xu, Zhi-Qin John
author_facet Wang, Zhiwei
Wang, Yunji
Zhang, Zhongwang
Zhou, Zhangchen
Jin, Hui
Hu, Tianyang
Sun, Jiacheng
Li, Zhenguo
Zhang, Yaoyu
Xu, Zhi-Qin John
contents Large language models have consistently struggled with complex reasoning tasks, such as mathematical problem-solving. Investigating the internal reasoning mechanisms of these models can help us design better model architectures and training strategies, ultimately enhancing their reasoning capability. In this study, we constructed a symbolic multi-step reasoning task to investigate the information propagation mechanisms in Transformer models when solving the task through direct answering and Chain-of-Thought (CoT) reasoning. We introduced the concept of buffer mechanism: the model stores various information in distinct buffers and selectively extracts it through the query-key matrix. We proposed a random matrix-based algorithm to enhance the model's reasoning ability. This algorithm introduces only 132 trainable parameters, yet leads to significant performance improvements on 7 multi-step reasoning datasets, including PrOntoQA, LogicAsker, and LogicInference. These findings provide new insights into understanding the large language models.
format Preprint
id arxiv_https___arxiv_org_abs_2405_15302
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Understanding the Language Model to Solve the Symbolic Multi-Step Reasoning Problem from the Perspective of Buffer Mechanism
Wang, Zhiwei
Wang, Yunji
Zhang, Zhongwang
Zhou, Zhangchen
Jin, Hui
Hu, Tianyang
Sun, Jiacheng
Li, Zhenguo
Zhang, Yaoyu
Xu, Zhi-Qin John
Artificial Intelligence
Computation and Language
Machine Learning
Large language models have consistently struggled with complex reasoning tasks, such as mathematical problem-solving. Investigating the internal reasoning mechanisms of these models can help us design better model architectures and training strategies, ultimately enhancing their reasoning capability. In this study, we constructed a symbolic multi-step reasoning task to investigate the information propagation mechanisms in Transformer models when solving the task through direct answering and Chain-of-Thought (CoT) reasoning. We introduced the concept of buffer mechanism: the model stores various information in distinct buffers and selectively extracts it through the query-key matrix. We proposed a random matrix-based algorithm to enhance the model's reasoning ability. This algorithm introduces only 132 trainable parameters, yet leads to significant performance improvements on 7 multi-step reasoning datasets, including PrOntoQA, LogicAsker, and LogicInference. These findings provide new insights into understanding the large language models.
title Understanding the Language Model to Solve the Symbolic Multi-Step Reasoning Problem from the Perspective of Buffer Mechanism
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2405.15302