A Primer on the Inner Workings of Transformer-based Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ferrando, Javier, Sarti, Gabriele, Bisazza, Arianna, Costa-jussà, Marta R.
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916435995918336
author Ferrando, Javier
Sarti, Gabriele
Bisazza, Arianna
Costa-jussà, Marta R.
author_facet Ferrando, Javier
Sarti, Gabriele
Bisazza, Arianna
Costa-jussà, Marta R.
contents The rapid progress of research aimed at interpreting the inner workings of advanced language models has highlighted a need for contextualizing the insights gained from years of work in this area. This primer provides a concise technical introduction to the current techniques used to interpret the inner workings of Transformer-based language models, focusing on the generative decoder-only architecture. We conclude by presenting a comprehensive overview of the known internal mechanisms implemented by these models, uncovering connections across popular approaches and active research directions in this area.
format Preprint
id arxiv_https___arxiv_org_abs_2405_00208
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Primer on the Inner Workings of Transformer-based Language Models
Ferrando, Javier
Sarti, Gabriele
Bisazza, Arianna
Costa-jussà, Marta R.
Computation and Language
The rapid progress of research aimed at interpreting the inner workings of advanced language models has highlighted a need for contextualizing the insights gained from years of work in this area. This primer provides a concise technical introduction to the current techniques used to interpret the inner workings of Transformer-based language models, focusing on the generative decoder-only architecture. We conclude by presenting a comprehensive overview of the known internal mechanisms implemented by these models, uncovering connections across popular approaches and active research directions in this area.
title A Primer on the Inner Workings of Transformer-based Language Models
topic Computation and Language
url https://arxiv.org/abs/2405.00208