Cascade Pipeline for Leading-Order Matrix Element Evaluation on AMD Versal AI Engine Arrays

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: López, P. Leguina, Villalba, C. Vico, Álvarez, F. Hervás, Arance, H. Gutiérrez, Folgueras, S., Fiorini, L., Valero, A., Menéndez, J. Fernández, Carrió, F., Oyanguren, A.
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914545627299840
author López, P. Leguina
Villalba, C. Vico
Álvarez, F. Hervás
Arance, H. Gutiérrez
Folgueras, S.
Fiorini, L.
Valero, A.
Menéndez, J. Fernández
Carrió, F.
Oyanguren, A.
author_facet López, P. Leguina
Villalba, C. Vico
Álvarez, F. Hervás
Arance, H. Gutiérrez
Folgueras, S.
Fiorini, L.
Valero, A.
Menéndez, J. Fernández
Carrió, F.
Oyanguren, A.
contents A major computational bottleneck in modern High Energy Physics event generators arises from the integration of the matrix element, which requires repeated evaluations at different phase-space points to cover all possible initial- and final-state configurations. As the Large Hadron Collider enters its High-Luminosity phase, the demand for energy-efficient acceleration is expected to exceed the limits of conventional CPU scaling, motivating the use of highly parallel computing platforms such as graphics processing units (GPUs). In this work, we present an alternative approach based on a cascade pipeline architecture for evaluating leading-order matrix elements of the \ggttg process on AMD Versal AI Engine (\aie) arrays. Due to the 16\,kB per-tile program memory constraint, the computation is decomposed into a five-stage pipeline, with stages communicating via a wavefunction-token protocol over the on-chip cascade interface. Mapping 80 independent pipelines onto the 400 \aie tiles of the VCK190 platform yields a projected throughput of $1.0\times10^6$ matrix element evaluations per second at 54.8\,W, corresponding to a $34\times$ speedup over a single CPU core and a $7.7\times$ improvement in energy efficiency. Numerical agreement with the \amcnlo double-precision reference is validated at the parts-per-million level in mean relative error.
format Preprint
id arxiv_https___arxiv_org_abs_2605_02481
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Cascade Pipeline for Leading-Order Matrix Element Evaluation on AMD Versal AI Engine Arrays
López, P. Leguina
Villalba, C. Vico
Álvarez, F. Hervás
Arance, H. Gutiérrez
Folgueras, S.
Fiorini, L.
Valero, A.
Menéndez, J. Fernández
Carrió, F.
Oyanguren, A.
High Energy Physics - Experiment
A major computational bottleneck in modern High Energy Physics event generators arises from the integration of the matrix element, which requires repeated evaluations at different phase-space points to cover all possible initial- and final-state configurations. As the Large Hadron Collider enters its High-Luminosity phase, the demand for energy-efficient acceleration is expected to exceed the limits of conventional CPU scaling, motivating the use of highly parallel computing platforms such as graphics processing units (GPUs). In this work, we present an alternative approach based on a cascade pipeline architecture for evaluating leading-order matrix elements of the \ggttg process on AMD Versal AI Engine (\aie) arrays. Due to the 16\,kB per-tile program memory constraint, the computation is decomposed into a five-stage pipeline, with stages communicating via a wavefunction-token protocol over the on-chip cascade interface. Mapping 80 independent pipelines onto the 400 \aie tiles of the VCK190 platform yields a projected throughput of $1.0\times10^6$ matrix element evaluations per second at 54.8\,W, corresponding to a $34\times$ speedup over a single CPU core and a $7.7\times$ improvement in energy efficiency. Numerical agreement with the \amcnlo double-precision reference is validated at the parts-per-million level in mean relative error.
title Cascade Pipeline for Leading-Order Matrix Element Evaluation on AMD Versal AI Engine Arrays
topic High Energy Physics - Experiment
url https://arxiv.org/abs/2605.02481