The emergence of sparse attention: impact of data distribution and benefits of repetition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zucchet, Nicolas, d'Angelo, Francesco, Lampinen, Andrew K., Chan, Stephanie C. Y.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912756577337344
author Zucchet, Nicolas
d'Angelo, Francesco
Lampinen, Andrew K.
Chan, Stephanie C. Y.
author_facet Zucchet, Nicolas
d'Angelo, Francesco
Lampinen, Andrew K.
Chan, Stephanie C. Y.
contents Emergence is a fascinating property of large language models and neural networks more broadly: as models scale and train for longer, they sometimes develop new abilities in sudden ways. Despite initial studies, we still lack a comprehensive understanding of how and when these abilities emerge. To address this gap, we study the emergence over training of sparse attention, a critical and frequently observed attention pattern in Transformers. By combining theoretical analysis of a toy model with empirical observations on small Transformers trained on a linear regression variant, we uncover the mechanics driving sparse attention emergence and reveal that emergence timing follows power laws based on task structure, architecture, and optimizer choice. We additionally find that repetition can greatly speed up emergence. Finally, we confirm these results on a well-studied in-context associative recall task. Our findings provide a simple, theoretically grounded framework for understanding how data distributions and model design influence the learning dynamics behind one form of emergence.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17863
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The emergence of sparse attention: impact of data distribution and benefits of repetition
Zucchet, Nicolas
d'Angelo, Francesco
Lampinen, Andrew K.
Chan, Stephanie C. Y.
Machine Learning
Neural and Evolutionary Computing
Emergence is a fascinating property of large language models and neural networks more broadly: as models scale and train for longer, they sometimes develop new abilities in sudden ways. Despite initial studies, we still lack a comprehensive understanding of how and when these abilities emerge. To address this gap, we study the emergence over training of sparse attention, a critical and frequently observed attention pattern in Transformers. By combining theoretical analysis of a toy model with empirical observations on small Transformers trained on a linear regression variant, we uncover the mechanics driving sparse attention emergence and reveal that emergence timing follows power laws based on task structure, architecture, and optimizer choice. We additionally find that repetition can greatly speed up emergence. Finally, we confirm these results on a well-studied in-context associative recall task. Our findings provide a simple, theoretically grounded framework for understanding how data distributions and model design influence the learning dynamics behind one form of emergence.
title The emergence of sparse attention: impact of data distribution and benefits of repetition
topic Machine Learning
Neural and Evolutionary Computing
url https://arxiv.org/abs/2505.17863