Associative memory inspires improvements for in-context learning using a novel attention residual stream architecture

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Burns, Thomas F, Fukai, Tomoki, Earls, Christopher J
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918113018118144
author Burns, Thomas F
Fukai, Tomoki
Earls, Christopher J
author_facet Burns, Thomas F
Fukai, Tomoki
Earls, Christopher J
contents Large language models (LLMs) demonstrate an impressive ability to utilise information within the context of their input sequences to appropriately respond to data unseen by the LLM during its training procedure. This ability is known as in-context learning (ICL). Humans and non-human animals demonstrate similar abilities, however their neural architectures differ substantially from LLMs. Despite this, a critical component within LLMs, the attention mechanism, resembles modern associative memory models, widely used in and influenced by the computational neuroscience community to model biological memory systems. Using this connection, we introduce an associative memory model capable of performing ICL. We use this as inspiration for a novel residual stream architecture which allows information to directly flow between attention heads. We test this architecture during training within a two-layer Transformer and show its ICL abilities manifest more quickly than without this modification. We then apply our architecture in small language models with 8 million and 1 billion parameters, focusing on attention head values, with results also indicating improved performance at these larger and more naturalistic scales.
format Preprint
id arxiv_https___arxiv_org_abs_2412_15113
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Associative memory inspires improvements for in-context learning using a novel attention residual stream architecture
Burns, Thomas F
Fukai, Tomoki
Earls, Christopher J
Neural and Evolutionary Computing
Artificial Intelligence
Computation and Language
92B20, 68T01, 68T37, 68T50
I.2; I.5; I.7; J.2; J.3
Large language models (LLMs) demonstrate an impressive ability to utilise information within the context of their input sequences to appropriately respond to data unseen by the LLM during its training procedure. This ability is known as in-context learning (ICL). Humans and non-human animals demonstrate similar abilities, however their neural architectures differ substantially from LLMs. Despite this, a critical component within LLMs, the attention mechanism, resembles modern associative memory models, widely used in and influenced by the computational neuroscience community to model biological memory systems. Using this connection, we introduce an associative memory model capable of performing ICL. We use this as inspiration for a novel residual stream architecture which allows information to directly flow between attention heads. We test this architecture during training within a two-layer Transformer and show its ICL abilities manifest more quickly than without this modification. We then apply our architecture in small language models with 8 million and 1 billion parameters, focusing on attention head values, with results also indicating improved performance at these larger and more naturalistic scales.
title Associative memory inspires improvements for in-context learning using a novel attention residual stream architecture
topic Neural and Evolutionary Computing
Artificial Intelligence
Computation and Language
92B20, 68T01, 68T37, 68T50
I.2; I.5; I.7; J.2; J.3
url https://arxiv.org/abs/2412.15113