States Hidden in Hidden States: LLMs Emerge Discrete State Representations Implicitly

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Junhao, Hu, Shengding, Liu, Zhiyuan, Sun, Maosong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913432640421888
author Chen, Junhao
Hu, Shengding
Liu, Zhiyuan
Sun, Maosong
author_facet Chen, Junhao
Hu, Shengding
Liu, Zhiyuan
Sun, Maosong
contents Large Language Models (LLMs) exhibit various emergent abilities. Among these abilities, some might reveal the internal working mechanisms of models. In this paper, we uncover a novel emergent capability in models: the intrinsic ability to perform extended sequences of calculations without relying on chain-of-thought step-by-step solutions. Remarkably, the most advanced models can directly output the results of two-digit number additions with lengths extending up to 15 addends. We hypothesize that the model emerges Implicit Discrete State Representations (IDSRs) within its hidden states and performs symbolic calculations internally. To test this hypothesis, we design a sequence of experiments that look into the hidden states. Specifically, we first confirm that IDSRs exist. Then, we provide interesting observations about the formation of IDSRs from layer, digit, and sequence perspectives. Finally, we confirm that models indeed use IDSRs to produce the final answers. However, we also discover that these state representations are far from lossless in current open-sourced models, leading to inaccuracies in their final performance. Our work presents a novel exploration of LLMs' symbolic calculation abilities and the underlying mechanisms.
format Preprint
id arxiv_https___arxiv_org_abs_2407_11421
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle States Hidden in Hidden States: LLMs Emerge Discrete State Representations Implicitly
Chen, Junhao
Hu, Shengding
Liu, Zhiyuan
Sun, Maosong
Computation and Language
Large Language Models (LLMs) exhibit various emergent abilities. Among these abilities, some might reveal the internal working mechanisms of models. In this paper, we uncover a novel emergent capability in models: the intrinsic ability to perform extended sequences of calculations without relying on chain-of-thought step-by-step solutions. Remarkably, the most advanced models can directly output the results of two-digit number additions with lengths extending up to 15 addends. We hypothesize that the model emerges Implicit Discrete State Representations (IDSRs) within its hidden states and performs symbolic calculations internally. To test this hypothesis, we design a sequence of experiments that look into the hidden states. Specifically, we first confirm that IDSRs exist. Then, we provide interesting observations about the formation of IDSRs from layer, digit, and sequence perspectives. Finally, we confirm that models indeed use IDSRs to produce the final answers. However, we also discover that these state representations are far from lossless in current open-sourced models, leading to inaccuracies in their final performance. Our work presents a novel exploration of LLMs' symbolic calculation abilities and the underlying mechanisms.
title States Hidden in Hidden States: LLMs Emerge Discrete State Representations Implicitly
topic Computation and Language
url https://arxiv.org/abs/2407.11421