Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining
Fuente:
arXiv
Guardado en:
| Autores principales: | , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866917273273368576 |
|---|---|
| author | Mitchell, Rupert Kersting, Kristian |
| author_facet | Mitchell, Rupert Kersting, Kristian |
| contents | Pretraining transformers on long sequences (entire code repositories, collections of related documents) is bottlenecked by quadratic attention costs. We present Multipole Semantic Attention (MuSe), which accelerates 64k-context pretraining by 36% while matching baseline loss, requiring no architectural changes. MuSe clusters queries and keys separately in representation space. This yields query-specific summaries that substantially outperform spatial blocking at matched sparsity, while also enabling drop-in compatibility with existing pretrained models; we validate on Llama 3.1-8B and 3.2-1B without retraining. We pretrain language models up to 1B parameters at 64k context on code and scientific documents, confirming that MuSe preserves quality and long-context utilization during training. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_10406 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining Mitchell, Rupert Kersting, Kristian Machine Learning 68W25, 68T50 (primary) 68W40, 68T07 (secondary) I.2.6; I.2.7 Pretraining transformers on long sequences (entire code repositories, collections of related documents) is bottlenecked by quadratic attention costs. We present Multipole Semantic Attention (MuSe), which accelerates 64k-context pretraining by 36% while matching baseline loss, requiring no architectural changes. MuSe clusters queries and keys separately in representation space. This yields query-specific summaries that substantially outperform spatial blocking at matched sparsity, while also enabling drop-in compatibility with existing pretrained models; we validate on Llama 3.1-8B and 3.2-1B without retraining. We pretrain language models up to 1B parameters at 64k context on code and scientific documents, confirming that MuSe preserves quality and long-context utilization during training. |
| title | Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining |
| topic | Machine Learning 68W25, 68T50 (primary) 68W40, 68T07 (secondary) I.2.6; I.2.7 |
| url | https://arxiv.org/abs/2509.10406 |