Integrated electro-optic attention nonlinearities for transformers

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Mickeler, Luis, Lion, Kai, Nardi, Alfonso, Kellner, Jost, Didier, Pierre, Shastri, Bhavin J., He, Niao, Grange, Rachel
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911582808702976
author Mickeler, Luis
Lion, Kai
Nardi, Alfonso
Kellner, Jost
Didier, Pierre
Shastri, Bhavin J.
He, Niao
Grange, Rachel
author_facet Mickeler, Luis
Lion, Kai
Nardi, Alfonso
Kellner, Jost
Didier, Pierre
Shastri, Bhavin J.
He, Niao
Grange, Rachel
contents Transformers have emerged as the dominant neural-network architecture, achieving state-of-the-art performance in language processing and computer vision. At the core of these models lies the attention mechanism, which requires a nonlinear, non-negative mapping using the Softmax function. However, although Softmax operations account for less than 1% of the total operation count, they can disproportionately bottleneck overall inference latency. Here, we use thin-film lithium niobate (TFLN) Mach-Zehnder modulators (MZMs) as analog nonlinear computational elements to drastically reduce the latency of nonlinear computations. We implement electro-optic alternatives to digital Softmax and Sigmoid, and evaluate their performance in Vision Transformers and Large Language Models. Our system maintains highly competitive accuracy, even under aggressive 4-bit input-output quantization of the analog units. We further characterize system noise at encoding speeds up to 10 GBaud and assess model robustness under various noise conditions. Our findings suggest that TFLN modulators can serve as nonlinear function units within hybrid co-packaged hardware, enabling high-speed and energy-efficient nonlinear computation.
format Preprint
id arxiv_https___arxiv_org_abs_2604_09512
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Integrated electro-optic attention nonlinearities for transformers
Mickeler, Luis
Lion, Kai
Nardi, Alfonso
Kellner, Jost
Didier, Pierre
Shastri, Bhavin J.
He, Niao
Grange, Rachel
Machine Learning
Optics
Transformers have emerged as the dominant neural-network architecture, achieving state-of-the-art performance in language processing and computer vision. At the core of these models lies the attention mechanism, which requires a nonlinear, non-negative mapping using the Softmax function. However, although Softmax operations account for less than 1% of the total operation count, they can disproportionately bottleneck overall inference latency. Here, we use thin-film lithium niobate (TFLN) Mach-Zehnder modulators (MZMs) as analog nonlinear computational elements to drastically reduce the latency of nonlinear computations. We implement electro-optic alternatives to digital Softmax and Sigmoid, and evaluate their performance in Vision Transformers and Large Language Models. Our system maintains highly competitive accuracy, even under aggressive 4-bit input-output quantization of the analog units. We further characterize system noise at encoding speeds up to 10 GBaud and assess model robustness under various noise conditions. Our findings suggest that TFLN modulators can serve as nonlinear function units within hybrid co-packaged hardware, enabling high-speed and energy-efficient nonlinear computation.
title Integrated electro-optic attention nonlinearities for transformers
topic Machine Learning
Optics
url https://arxiv.org/abs/2604.09512