Iterative Refinement with Low-Precision Posits
Fuente:
arXiv
Guardado en:
| Autores principales: | , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866909298033950720 |
|---|---|
| author | Quinlan, James Omtzigt, E. Theodore L. |
| author_facet | Quinlan, James Omtzigt, E. Theodore L. |
| contents | This research investigates using a mixed-precision iterative refinement method using posit numbers instead of the standard IEEE floating-point format. The method is applied to solve a general linear system represented by the equation $Ax = b$, where $A$ is a large sparse matrix. Various scaling techniques, such as row and column equilibration, map the matrix entries to higher-density regions of machine numbers before performing the $O(n^3)$ factorization operation. Low-precision LU factorization followed by forward/backward substitution provides an initial estimate. The results demonstrate that a 16-bit posit configuration combined with equilibration produces accuracy comparable to IEEE half-precision (fp16), indicating a potential for achieving a balance between efficiency and accuracy. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2408_13400 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Iterative Refinement with Low-Precision Posits Quinlan, James Omtzigt, E. Theodore L. Numerical Analysis 65F05 (Primary) 65F50 (Secondary) This research investigates using a mixed-precision iterative refinement method using posit numbers instead of the standard IEEE floating-point format. The method is applied to solve a general linear system represented by the equation $Ax = b$, where $A$ is a large sparse matrix. Various scaling techniques, such as row and column equilibration, map the matrix entries to higher-density regions of machine numbers before performing the $O(n^3)$ factorization operation. Low-precision LU factorization followed by forward/backward substitution provides an initial estimate. The results demonstrate that a 16-bit posit configuration combined with equilibration produces accuracy comparable to IEEE half-precision (fp16), indicating a potential for achieving a balance between efficiency and accuracy. |
| title | Iterative Refinement with Low-Precision Posits |
| topic | Numerical Analysis 65F05 (Primary) 65F50 (Secondary) |
| url | https://arxiv.org/abs/2408.13400 |