Characterizing stable regions in the residual stream of LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Janiak, Jett, Karwowski, Jacek, Mangat, Chatrik Singh, Giglemiani, Giorgi, Petrova, Nora, Heimersheim, Stefan
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913579644485632
author Janiak, Jett
Karwowski, Jacek
Mangat, Chatrik Singh
Giglemiani, Giorgi
Petrova, Nora
Heimersheim, Stefan
author_facet Janiak, Jett
Karwowski, Jacek
Mangat, Chatrik Singh
Giglemiani, Giorgi
Petrova, Nora
Heimersheim, Stefan
contents We identify stable regions in the residual stream of Transformers, where the model's output remains insensitive to small activation changes, but exhibits high sensitivity at region boundaries. These regions emerge during training and become more defined as training progresses or model size increases. The regions appear to be much larger than previously studied polytopes. Our analysis suggests that these stable regions align with semantic distinctions, where similar prompts cluster within regions, and activations from the same region lead to similar next token predictions. This work provides a promising research direction for understanding the complexity of neural networks, shedding light on training dynamics, and advancing interpretability.
format Preprint
id arxiv_https___arxiv_org_abs_2409_17113
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Characterizing stable regions in the residual stream of LLMs
Janiak, Jett
Karwowski, Jacek
Mangat, Chatrik Singh
Giglemiani, Giorgi
Petrova, Nora
Heimersheim, Stefan
Machine Learning
We identify stable regions in the residual stream of Transformers, where the model's output remains insensitive to small activation changes, but exhibits high sensitivity at region boundaries. These regions emerge during training and become more defined as training progresses or model size increases. The regions appear to be much larger than previously studied polytopes. Our analysis suggests that these stable regions align with semantic distinctions, where similar prompts cluster within regions, and activations from the same region lead to similar next token predictions. This work provides a promising research direction for understanding the complexity of neural networks, shedding light on training dynamics, and advancing interpretability.
title Characterizing stable regions in the residual stream of LLMs
topic Machine Learning
url https://arxiv.org/abs/2409.17113