Performance and Numerical Aspects of Decompositional Factorizations with FP64 Floating-Point Emulation in INT8
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866911181993672704 |
|---|---|
| author | Luszczek, Piotr Gadepally, Vijay Anderson, LaToya Arcand, William Bestor, David Bergeron, William Bonn, Alex Burrill, Daniel J. Byun, Chansup Houle, Michael Hubbell, Matthew Jananthan, Hayden Jones, Michael Michaleas, Peter Morales, Guillermo Mullen, Julia Prout, Andrew Reuther, Albert Rosa, Antonio Yee, Charles Kepner, Jeremy |
| author_facet | Luszczek, Piotr Gadepally, Vijay Anderson, LaToya Arcand, William Bestor, David Bergeron, William Bonn, Alex Burrill, Daniel J. Byun, Chansup Houle, Michael Hubbell, Matthew Jananthan, Hayden Jones, Michael Michaleas, Peter Morales, Guillermo Mullen, Julia Prout, Andrew Reuther, Albert Rosa, Antonio Yee, Charles Kepner, Jeremy |
| contents | Mixing precisions for performance has been an ongoing trend as the modern hardware accelerators started including new, and mostly lower-precision, data formats. The advantage of using them is a great potential of performance gain and energy savings. The disadvantage are the numerical issues not present in the standard-mandated floating-point formats. Split integer emulation of FP64 takes this to an extreme with the computation performed only by fixed-point tensor core units. We present the new issues the emulation faces for practical cases involving dense linear solver. We show extensive numerical tests indicating the effect of extended numerical range of matrix entries. We also scaled the input sizes to study the performance and numerical profiles on the NVIDIA Hopper GPUs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_23565 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Performance and Numerical Aspects of Decompositional Factorizations with FP64 Floating-Point Emulation in INT8 Luszczek, Piotr Gadepally, Vijay Anderson, LaToya Arcand, William Bestor, David Bergeron, William Bonn, Alex Burrill, Daniel J. Byun, Chansup Houle, Michael Hubbell, Matthew Jananthan, Hayden Jones, Michael Michaleas, Peter Morales, Guillermo Mullen, Julia Prout, Andrew Reuther, Albert Rosa, Antonio Yee, Charles Kepner, Jeremy Numerical Analysis Mixing precisions for performance has been an ongoing trend as the modern hardware accelerators started including new, and mostly lower-precision, data formats. The advantage of using them is a great potential of performance gain and energy savings. The disadvantage are the numerical issues not present in the standard-mandated floating-point formats. Split integer emulation of FP64 takes this to an extreme with the computation performed only by fixed-point tensor core units. We present the new issues the emulation faces for practical cases involving dense linear solver. We show extensive numerical tests indicating the effect of extended numerical range of matrix entries. We also scaled the input sizes to study the performance and numerical profiles on the NVIDIA Hopper GPUs. |
| title | Performance and Numerical Aspects of Decompositional Factorizations with FP64 Floating-Point Emulation in INT8 |
| topic | Numerical Analysis |
| url | https://arxiv.org/abs/2509.23565 |