Exploring Information Processing in Large Language Models: Insights from Information Bottleneck Theory

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yang, Zhou, Qi, Zhengyu, Ren, Zhaochun, Jia, Zhikai, Sun, Haizhou, Zhu, Xiaofei, Liao, Xiangwen
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929659528085504
author Yang, Zhou
Qi, Zhengyu
Ren, Zhaochun
Jia, Zhikai
Sun, Haizhou
Zhu, Xiaofei
Liao, Xiangwen
author_facet Yang, Zhou
Qi, Zhengyu
Ren, Zhaochun
Jia, Zhikai
Sun, Haizhou
Zhu, Xiaofei
Liao, Xiangwen
contents Large Language Models (LLMs) have demonstrated remarkable performance across a wide range of tasks by understanding input information and predicting corresponding outputs. However, the internal mechanisms by which LLMs comprehend input and make effective predictions remain poorly understood. In this paper, we explore the working mechanism of LLMs in information processing from the perspective of Information Bottleneck Theory. We propose a non-training construction strategy to define a task space and identify the following key findings: (1) LLMs compress input information into specific task spaces (e.g., sentiment space, topic space) to facilitate task understanding; (2) they then extract and utilize relevant information from the task space at critical moments to generate accurate predictions. Based on these insights, we introduce two novel approaches: an Information Compression-based Context Learning (IC-ICL) and a Task-Space-guided Fine-Tuning (TS-FT). IC-ICL enhances reasoning performance and inference efficiency by compressing retrieved example information into the task space. TS-FT employs a space-guided loss to fine-tune LLMs, encouraging the learning of more effective compression and selection mechanisms. Experiments across multiple datasets validate the effectiveness of task space construction. Additionally, IC-ICL not only improves performance but also accelerates inference speed by over 40\%, while TS-FT achieves superior results with a minimal strategy adjustment.
format Preprint
id arxiv_https___arxiv_org_abs_2501_00999
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring Information Processing in Large Language Models: Insights from Information Bottleneck Theory
Yang, Zhou
Qi, Zhengyu
Ren, Zhaochun
Jia, Zhikai
Sun, Haizhou
Zhu, Xiaofei
Liao, Xiangwen
Computation and Language
Artificial Intelligence
Large Language Models (LLMs) have demonstrated remarkable performance across a wide range of tasks by understanding input information and predicting corresponding outputs. However, the internal mechanisms by which LLMs comprehend input and make effective predictions remain poorly understood. In this paper, we explore the working mechanism of LLMs in information processing from the perspective of Information Bottleneck Theory. We propose a non-training construction strategy to define a task space and identify the following key findings: (1) LLMs compress input information into specific task spaces (e.g., sentiment space, topic space) to facilitate task understanding; (2) they then extract and utilize relevant information from the task space at critical moments to generate accurate predictions. Based on these insights, we introduce two novel approaches: an Information Compression-based Context Learning (IC-ICL) and a Task-Space-guided Fine-Tuning (TS-FT). IC-ICL enhances reasoning performance and inference efficiency by compressing retrieved example information into the task space. TS-FT employs a space-guided loss to fine-tune LLMs, encouraging the learning of more effective compression and selection mechanisms. Experiments across multiple datasets validate the effectiveness of task space construction. Additionally, IC-ICL not only improves performance but also accelerates inference speed by over 40\%, while TS-FT achieves superior results with a minimal strategy adjustment.
title Exploring Information Processing in Large Language Models: Insights from Information Bottleneck Theory
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2501.00999