High-Performance Pipelined NTT Accelerators with Homogeneous Digit-Serial Modulo Arithmetic
Fuente:
arXiv
Saved in:
| Main Authors: | Alexakis, George, Schoinianakis, Dimitrios, Dimitrakopoulos, Giorgos |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Implementation of RISC-V Vector Permutation Instructions
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
Register Dispersion: Reducing the Footprint of the Vector Register File in Vector Engines of Low-Cost RISC-V CPUs
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
H-FA: A Hybrid Floating-Point and Logarithmic Approach to Hardware Accelerated FlashAttention
by: Alexandridis, Kosmas, et al.
Published: (2025)
by: Alexandridis, Kosmas, et al.
Published: (2025)
Online Alignment and Addition in Multi-Term Floating-Point Adders
by: Alexandridis, Kosmas, et al.
Published: (2024)
by: Alexandridis, Kosmas, et al.
Published: (2024)
DSLR-CNN: Efficient CNN Acceleration using Digit-Serial Left-to-Right Arithmetic
by: Nisar, Malik Zohaib, et al.
Published: (2025)
by: Nisar, Malik Zohaib, et al.
Published: (2025)
Error Checking for Sparse Systolic Tensor Arrays
by: Peltekis, Christodoulos, et al.
Published: (2024)
by: Peltekis, Christodoulos, et al.
Published: (2024)
GCN-ABFT: Low-Cost Online Error Checking for Graph Convolutional Networks
by: Peltekis, Christodoulos, et al.
Published: (2024)
by: Peltekis, Christodoulos, et al.
Published: (2024)
Hermes: A Unified High-Performance NTT Architecture with Hybrid Dataflow
by: Gu, Hang, et al.
Published: (2026)
by: Gu, Hang, et al.
Published: (2026)
Floating-Point Multiply-Add with Approximate Normalization for Low-Cost Matrix Engines
by: Alexandridis, Kosmas, et al.
Published: (2024)
by: Alexandridis, Kosmas, et al.
Published: (2024)
DeMM: A Decoupled Matrix Multiplication Engine Supporting Relaxed Structured Sparsity
by: Peltekis, Christodoulos, et al.
Published: (2024)
by: Peltekis, Christodoulos, et al.
Published: (2024)
Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
Periodic Online Testing for Sparse Systolic Tensor Arrays
by: Peltekis, Christodoulos, et al.
Published: (2025)
by: Peltekis, Christodoulos, et al.
Published: (2025)
Low-Cost FlashAttention with Fused Exponential and Multiplication Hardware Operators
by: Alexandridis, Kosmas, et al.
Published: (2025)
by: Alexandridis, Kosmas, et al.
Published: (2025)
Reusing Softmax Hardware Unit for GELU Computation in Transformers
by: Peltekis, Christodoulos, et al.
Published: (2024)
by: Peltekis, Christodoulos, et al.
Published: (2024)
Optimizing Structured-Sparse Matrix Multiplication in RISC-V Vector Processors
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
HF-NTT: Hazard-Free Dataflow Accelerator for Number Theoretic Transform
by: Meng, Xiangchen, et al.
Published: (2024)
by: Meng, Xiangchen, et al.
Published: (2024)
FLASH-D: FlashAttention with Hidden Softmax Division
by: Alexandridis, Kosmas, et al.
Published: (2025)
by: Alexandridis, Kosmas, et al.
Published: (2025)
@NTT: Algorithm-Targeted NTT hardware acceleration via Design-Time Constant Optimization
by: Nabeel, Mohammed, et al.
Published: (2026)
by: Nabeel, Mohammed, et al.
Published: (2026)
Modulo-$(2^{2n}+1)$ Arithmetic via Two Parallel n-bit Residue Channels
by: Jaberipur, Ghassem, et al.
Published: (2024)
by: Jaberipur, Ghassem, et al.
Published: (2024)
A Bespoke Design Approach to Low-Power Printed Microprocessors for Machine Learning Applications
by: Chaidos, Panagiotis, et al.
Published: (2025)
by: Chaidos, Panagiotis, et al.
Published: (2025)
MaRVIn: A Cross-Layer Mixed-Precision RISC-V Framework for DNN Inference, from ISA Extension to Hardware Acceleration
by: Armeniakos, Giorgos, et al.
Published: (2025)
by: Armeniakos, Giorgos, et al.
Published: (2025)
In-Pipeline Integration of Digital In-Memory-Computing into RISC-V Vector Architecture to Accelerate Deep Learning
by: Spagnolo, Tommaso, et al.
Published: (2026)
by: Spagnolo, Tommaso, et al.
Published: (2026)
bitSMM: A bit-Serial Matrix Multiplication Accelerator
by: Antunes, Pedro, et al.
Published: (2026)
by: Antunes, Pedro, et al.
Published: (2026)
PaReNTT: Low-Latency Parallel Residue Number System and NTT-Based Long Polynomial Modular Multiplication for Homomorphic Encryption
by: Tan, Weihang, et al.
Published: (2023)
by: Tan, Weihang, et al.
Published: (2023)
MAx-DNN: Multi-Level Arithmetic Approximation for Energy-Efficient DNN Hardware Accelerators
by: Leon, Vasileios, et al.
Published: (2025)
by: Leon, Vasileios, et al.
Published: (2025)
Towards Employing FPGA and ASIP Acceleration to Enable Onboard AI/ML in Space Applications
by: Leon, Vasileios, et al.
Published: (2025)
by: Leon, Vasileios, et al.
Published: (2025)
Combining Power and Arithmetic Optimization via Datapath Rewriting
by: Coward, Samuel, et al.
Published: (2024)
by: Coward, Samuel, et al.
Published: (2024)
TATAA: Programmable Mixed-Precision Transformer Acceleration with a Transformable Arithmetic Architecture
by: Wu, Jiajun, et al.
Published: (2024)
by: Wu, Jiajun, et al.
Published: (2024)
AccelSync: Verifying Synchronization Coverage in Accelerator Pipeline Programs
by: An, Hangcheng, et al.
Published: (2026)
by: An, Hangcheng, et al.
Published: (2026)
From Circuits to SoC Processors: Arithmetic Approximation Techniques & Embedded Computing Methodologies for DSP Acceleration
by: Leon, Vasileios
Published: (2023)
by: Leon, Vasileios
Published: (2023)
Development of High-Performance DSP Algorithms on the European Rad-Hard NG-ULTRA SoC FPGA
by: Leon, Vasileios, et al.
Published: (2024)
by: Leon, Vasileios, et al.
Published: (2024)
Trojan-Resilient NTT: Protecting Against Control Flow and Timing Faults on Reconfigurable Platforms
by: Paul, Rourab, et al.
Published: (2026)
by: Paul, Rourab, et al.
Published: (2026)
RPCAcc: A High-Performance and Reconfigurable PCIe-attached RPC Accelerator
by: Zhang, Jie, et al.
Published: (2024)
by: Zhang, Jie, et al.
Published: (2024)
HFRWKV: A High-Performance Fully On-Chip Hardware Accelerator for RWKV
by: Shijie, Liu, et al.
Published: (2026)
by: Shijie, Liu, et al.
Published: (2026)
Combining Fault Tolerance Techniques and COTS SoC Accelerators for Payload Processing in Space
by: Leon, Vasileios, et al.
Published: (2025)
by: Leon, Vasileios, et al.
Published: (2025)
Edge GPU Aware Multiple AI Model Pipeline for Accelerated MRI Reconstruction and Analysis
by: Majeed, Ashiyana Abdul, et al.
Published: (2025)
by: Majeed, Ashiyana Abdul, et al.
Published: (2025)
Scope: A Scalable Merged Pipeline Framework for Multi-Chip-Module NN Accelerators
by: Huang, Zongle, et al.
Published: (2026)
by: Huang, Zongle, et al.
Published: (2026)
PHAROS: Pipelined Heterogeneous Accelerators for Real-time Safety-critical Systems With Deadline Compliance
by: Ji, Shixin, et al.
Published: (2026)
by: Ji, Shixin, et al.
Published: (2026)
SCE-NTT: A Hardware Accelerator for Number Theoretic Transform Using Superconductor Electronics
by: Razmkhah, Sasan, et al.
Published: (2025)
by: Razmkhah, Sasan, et al.
Published: (2025)
Mixed-precision Neural Networks on RISC-V Cores: ISA extensions for Multi-Pumped Soft SIMD Operations
by: Armeniakos, Giorgos, et al.
Published: (2024)
by: Armeniakos, Giorgos, et al.
Published: (2024)
Similar Items
-
Efficient Implementation of RISC-V Vector Permutation Instructions
by: Titopoulos, Vasileios, et al.
Published: (2025) -
Register Dispersion: Reducing the Footprint of the Vector Register File in Vector Engines of Low-Cost RISC-V CPUs
by: Titopoulos, Vasileios, et al.
Published: (2025) -
H-FA: A Hybrid Floating-Point and Logarithmic Approach to Hardware Accelerated FlashAttention
by: Alexandridis, Kosmas, et al.
Published: (2025) -
Online Alignment and Addition in Multi-Term Floating-Point Adders
by: Alexandridis, Kosmas, et al.
Published: (2024) -
DSLR-CNN: Efficient CNN Acceleration using Digit-Serial Left-to-Right Arithmetic
by: Nisar, Malik Zohaib, et al.
Published: (2025)