Iterative Refinement with Low-Precision Posits

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Quinlan, James, Omtzigt, E. Theodore L.
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909298033950720
author Quinlan, James
Omtzigt, E. Theodore L.
author_facet Quinlan, James
Omtzigt, E. Theodore L.
contents This research investigates using a mixed-precision iterative refinement method using posit numbers instead of the standard IEEE floating-point format. The method is applied to solve a general linear system represented by the equation $Ax = b$, where $A$ is a large sparse matrix. Various scaling techniques, such as row and column equilibration, map the matrix entries to higher-density regions of machine numbers before performing the $O(n^3)$ factorization operation. Low-precision LU factorization followed by forward/backward substitution provides an initial estimate. The results demonstrate that a 16-bit posit configuration combined with equilibration produces accuracy comparable to IEEE half-precision (fp16), indicating a potential for achieving a balance between efficiency and accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2408_13400
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Iterative Refinement with Low-Precision Posits
Quinlan, James
Omtzigt, E. Theodore L.
Numerical Analysis
65F05 (Primary) 65F50 (Secondary)
This research investigates using a mixed-precision iterative refinement method using posit numbers instead of the standard IEEE floating-point format. The method is applied to solve a general linear system represented by the equation $Ax = b$, where $A$ is a large sparse matrix. Various scaling techniques, such as row and column equilibration, map the matrix entries to higher-density regions of machine numbers before performing the $O(n^3)$ factorization operation. Low-precision LU factorization followed by forward/backward substitution provides an initial estimate. The results demonstrate that a 16-bit posit configuration combined with equilibration produces accuracy comparable to IEEE half-precision (fp16), indicating a potential for achieving a balance between efficiency and accuracy.
title Iterative Refinement with Low-Precision Posits
topic Numerical Analysis
65F05 (Primary) 65F50 (Secondary)
url https://arxiv.org/abs/2408.13400