Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Mitchell, Rupert, Kersting, Kristian
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917273273368576
author Mitchell, Rupert
Kersting, Kristian
author_facet Mitchell, Rupert
Kersting, Kristian
contents Pretraining transformers on long sequences (entire code repositories, collections of related documents) is bottlenecked by quadratic attention costs. We present Multipole Semantic Attention (MuSe), which accelerates 64k-context pretraining by 36% while matching baseline loss, requiring no architectural changes. MuSe clusters queries and keys separately in representation space. This yields query-specific summaries that substantially outperform spatial blocking at matched sparsity, while also enabling drop-in compatibility with existing pretrained models; we validate on Llama 3.1-8B and 3.2-1B without retraining. We pretrain language models up to 1B parameters at 64k context on code and scientific documents, confirming that MuSe preserves quality and long-context utilization during training.
format Preprint
id arxiv_https___arxiv_org_abs_2509_10406
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining
Mitchell, Rupert
Kersting, Kristian
Machine Learning
68W25, 68T50 (primary) 68W40, 68T07 (secondary)
I.2.6; I.2.7
Pretraining transformers on long sequences (entire code repositories, collections of related documents) is bottlenecked by quadratic attention costs. We present Multipole Semantic Attention (MuSe), which accelerates 64k-context pretraining by 36% while matching baseline loss, requiring no architectural changes. MuSe clusters queries and keys separately in representation space. This yields query-specific summaries that substantially outperform spatial blocking at matched sparsity, while also enabling drop-in compatibility with existing pretrained models; we validate on Llama 3.1-8B and 3.2-1B without retraining. We pretrain language models up to 1B parameters at 64k context on code and scientific documents, confirming that MuSe preserves quality and long-context utilization during training.
title Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining
topic Machine Learning
68W25, 68T50 (primary) 68W40, 68T07 (secondary)
I.2.6; I.2.7
url https://arxiv.org/abs/2509.10406