Benign Overfitting in Token Selection of Attention Mechanism

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sakamoto, Keitaro, Sato, Issei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912380400697344
author Sakamoto, Keitaro
Sato, Issei
author_facet Sakamoto, Keitaro
Sato, Issei
contents Attention mechanism is a fundamental component of the transformer model and plays a significant role in its success. However, the theoretical understanding of how attention learns to select tokens is still an emerging area of research. In this work, we study the training dynamics and generalization ability of the attention mechanism under classification problems with label noise. We show that, with the characterization of signal-to-noise ratio (SNR), the token selection of attention mechanism achieves benign overfitting, i.e., maintaining high generalization performance despite fitting label noise. Our work also demonstrates an interesting delayed acquisition of generalization after an initial phase of overfitting. Finally, we provide experiments to support our theoretical analysis using both synthetic and real-world datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2409_17625
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Benign Overfitting in Token Selection of Attention Mechanism
Sakamoto, Keitaro
Sato, Issei
Machine Learning
Attention mechanism is a fundamental component of the transformer model and plays a significant role in its success. However, the theoretical understanding of how attention learns to select tokens is still an emerging area of research. In this work, we study the training dynamics and generalization ability of the attention mechanism under classification problems with label noise. We show that, with the characterization of signal-to-noise ratio (SNR), the token selection of attention mechanism achieves benign overfitting, i.e., maintaining high generalization performance despite fitting label noise. Our work also demonstrates an interesting delayed acquisition of generalization after an initial phase of overfitting. Finally, we provide experiments to support our theoretical analysis using both synthetic and real-world datasets.
title Benign Overfitting in Token Selection of Attention Mechanism
topic Machine Learning
url https://arxiv.org/abs/2409.17625