Leveraging SIMD for Accelerating Large-number Arithmetic

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Das, Subhrajit, Bichhawat, Abhishek, Patel, Yuvraj
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917431626170368
author Das, Subhrajit
Bichhawat, Abhishek
Patel, Yuvraj
author_facet Das, Subhrajit
Bichhawat, Abhishek
Patel, Yuvraj
contents Large-number arithmetic, widely used in scientific computing and cryptography, has seen limited adoption of single instruction, multiple data (SIMD) parallelism on modern CPUs due to the inherent dependencies in traditional algorithms. We present DigitsOnTurbo (DoT), which restructures the computation around independent, data-parallel operations, rather than vectorizing the standard algorithms, thereby leveraging the benefits provided by SIMD. Over prior SIMD implementations, DoT achieves up to 1.85x speedups for addition and subtraction, and 2.3x for multiplication. When integrated into state-of-the-art libraries, DoT yields up to 4x speedup for addition and subtraction, and up to 2x speedup for multiplication, cascading into end-to-end throughput gains of up to 19.3% for scientific computations, and up to 7.9% latency and 5.9% throughput improvements on cryptographic implementations.
format Preprint
id arxiv_https___arxiv_org_abs_2604_21566
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Leveraging SIMD for Accelerating Large-number Arithmetic
Das, Subhrajit
Bichhawat, Abhishek
Patel, Yuvraj
Distributed, Parallel, and Cluster Computing
Hardware Architecture
Large-number arithmetic, widely used in scientific computing and cryptography, has seen limited adoption of single instruction, multiple data (SIMD) parallelism on modern CPUs due to the inherent dependencies in traditional algorithms. We present DigitsOnTurbo (DoT), which restructures the computation around independent, data-parallel operations, rather than vectorizing the standard algorithms, thereby leveraging the benefits provided by SIMD. Over prior SIMD implementations, DoT achieves up to 1.85x speedups for addition and subtraction, and 2.3x for multiplication. When integrated into state-of-the-art libraries, DoT yields up to 4x speedup for addition and subtraction, and up to 2x speedup for multiplication, cascading into end-to-end throughput gains of up to 19.3% for scientific computations, and up to 7.9% latency and 5.9% throughput improvements on cryptographic implementations.
title Leveraging SIMD for Accelerating Large-number Arithmetic
topic Distributed, Parallel, and Cluster Computing
Hardware Architecture
url https://arxiv.org/abs/2604.21566