Quantifying Semantic Emergence in Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Hang, Yang, Xinyu, Zhu, Jiaying, Wang, Wenya
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910750333730816
author Chen, Hang
Yang, Xinyu
Zhu, Jiaying
Wang, Wenya
author_facet Chen, Hang
Yang, Xinyu
Zhu, Jiaying
Wang, Wenya
contents Large language models (LLMs) are widely recognized for their exceptional capacity to capture semantics meaning. Yet, there remains no established metric to quantify this capability. In this work, we introduce a quantitative metric, Information Emergence (IE), designed to measure LLMs' ability to extract semantics from input tokens. We formalize ``semantics'' as the meaningful information abstracted from a sequence of tokens and quantify this by comparing the entropy reduction observed for a sequence of tokens (macro-level) and individual tokens (micro-level). To achieve this, we design a lightweight estimator to compute the mutual information at each transformer layer, which is agnostic to different tasks and language model architectures. We apply IE in both synthetic in-context learning (ICL) scenarios and natural sentence contexts. Experiments demonstrate informativeness and patterns about semantics. While some of these patterns confirm the conventional prior linguistic knowledge, the rest are relatively unexpected, which may provide new insights.
format Preprint
id arxiv_https___arxiv_org_abs_2405_12617
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Quantifying Semantic Emergence in Language Models
Chen, Hang
Yang, Xinyu
Zhu, Jiaying
Wang, Wenya
Computation and Language
Artificial Intelligence
Large language models (LLMs) are widely recognized for their exceptional capacity to capture semantics meaning. Yet, there remains no established metric to quantify this capability. In this work, we introduce a quantitative metric, Information Emergence (IE), designed to measure LLMs' ability to extract semantics from input tokens. We formalize ``semantics'' as the meaningful information abstracted from a sequence of tokens and quantify this by comparing the entropy reduction observed for a sequence of tokens (macro-level) and individual tokens (micro-level). To achieve this, we design a lightweight estimator to compute the mutual information at each transformer layer, which is agnostic to different tasks and language model architectures. We apply IE in both synthetic in-context learning (ICL) scenarios and natural sentence contexts. Experiments demonstrate informativeness and patterns about semantics. While some of these patterns confirm the conventional prior linguistic knowledge, the rest are relatively unexpected, which may provide new insights.
title Quantifying Semantic Emergence in Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2405.12617