Synthesis of signal processing algorithms with constraints on minimal parallelism and memory space
Fuente:
arXiv
Saved in:
| Main Author: | Salishev, Sergey |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Regular mixed-radix DFT matrix factorization for in-place FFT accelerators
by: Salishev, Sergey
Published: (2025)
by: Salishev, Sergey
Published: (2025)
PackSELL: A Sparse Matrix Format for Precision-Agnostic High-Performance SpMV
by: Suzuki, Kengo, et al.
Published: (2026)
by: Suzuki, Kengo, et al.
Published: (2026)
Wattchmen: Watching the Wattchers -- High Fidelity, Flexible GPU Energy Modeling
by: Tran, Brandon, et al.
Published: (2026)
by: Tran, Brandon, et al.
Published: (2026)
Exploiting nested task-parallelism in the $\mathcal{H}-LU$ factorization
by: Carratalá-Sáez, Rocío, et al.
Published: (2019)
by: Carratalá-Sáez, Rocío, et al.
Published: (2019)
NCCL EP: Towards a Unified Expert Parallel Communication API for NCCL
by: Goldman, Amos, et al.
Published: (2026)
by: Goldman, Amos, et al.
Published: (2026)
Data Scheduling Algorithm for Scalable and Efficient IoT Sensing in Cloud Computing
by: Mohammad, Noor Islam S.
Published: (2025)
by: Mohammad, Noor Islam S.
Published: (2025)
GPU-Initiated Networking for NCCL
by: Hamidouche, Khaled, et al.
Published: (2025)
by: Hamidouche, Khaled, et al.
Published: (2025)
A scalable high-order multigrid-FFT Poisson solver for unbounded domains on adaptive multiresolution grids
by: Poncelet, Gilles, et al.
Published: (2025)
by: Poncelet, Gilles, et al.
Published: (2025)
Distributed Tomographic Reconstruction with Quantization
by: Miao, Runxuan, et al.
Published: (2024)
by: Miao, Runxuan, et al.
Published: (2024)
Parallelization Strategies for the Randomized Kaczmarz Algorithm on Large-Scale Dense Systems
by: Ferreira, Inês, et al.
Published: (2024)
by: Ferreira, Inês, et al.
Published: (2024)
Precision-Aware Iterative Algorithms Based on Group-Shared Exponents of Floating-Point Numbers
by: Gao, Jianhua, et al.
Published: (2024)
by: Gao, Jianhua, et al.
Published: (2024)
Efficient Parallel Scheduling for Sparse Triangular Solvers
by: Böhnlein, Toni, et al.
Published: (2025)
by: Böhnlein, Toni, et al.
Published: (2025)
Floating Point Compression of Hierarchical Matrix Formats and its Impact on Matrix-Vector Multiplication
by: Kriemann, Ronald
Published: (2024)
by: Kriemann, Ronald
Published: (2024)
RapidStream IR: Infrastructure for FPGA High-Level Physical Synthesis
by: Lau, Jason, et al.
Published: (2024)
by: Lau, Jason, et al.
Published: (2024)
A new Dune grid for scalable dynamic adaptivity based on the p4est software library
by: Burstedde, Carsten, et al.
Published: (2025)
by: Burstedde, Carsten, et al.
Published: (2025)
Method for determining the acceleration of a parallel specialised computer system based on Amdahl's law
by: Filipchenko, Aleksandr S.
Published: (2024)
by: Filipchenko, Aleksandr S.
Published: (2024)
Serial Parallel Reliability Redundancy Allocation Optimization for Energy Efficient and Fault Tolerant Cloud Computing
by: Krishna, Gutha Jaya
Published: (2024)
by: Krishna, Gutha Jaya
Published: (2024)
Distributed Hybrid Sketching for $\ell_2$-Embeddings
by: Charalambides, Neophytos, et al.
Published: (2024)
by: Charalambides, Neophytos, et al.
Published: (2024)
Accelerating State-Vector Quantum Simulation on Integrated GPUs via Cache Locality Optimization: A Cross-Architecture Evaluation
by: Thomaz, Gabriel Fernandes, et al.
Published: (2026)
by: Thomaz, Gabriel Fernandes, et al.
Published: (2026)
A multigrid reduction framework for domains with symmetries
by: Alsalti-Baldellou, Àdel, et al.
Published: (2024)
by: Alsalti-Baldellou, Àdel, et al.
Published: (2024)
The Performance of Low-Synchronization Variants of Reorthogonalized Block Classical Gram--Schmidt
by: Carson, Erin, et al.
Published: (2025)
by: Carson, Erin, et al.
Published: (2025)
Efficient Hardware Accelerator Based on Medium Granularity Dataflow for SpTRSV
by: Chen, Qian, et al.
Published: (2024)
by: Chen, Qian, et al.
Published: (2024)
A Systematic Literature Survey of Sparse Matrix-Vector Multiplication
by: Gao, Jianhua, et al.
Published: (2024)
by: Gao, Jianhua, et al.
Published: (2024)
Mixed-Precision Performance Portability of FFT-Based GPU-Accelerated Algorithms for Block-Triangular Toeplitz Matrices
by: Venkat, Sreeram, et al.
Published: (2025)
by: Venkat, Sreeram, et al.
Published: (2025)
On a randomized small-block Lanczos method for large-scale null space computations
by: Kressner, Daniel, et al.
Published: (2024)
by: Kressner, Daniel, et al.
Published: (2024)
Evaluation of computational and energy performance in matrix multiplication algorithms on CPU and GPU using MKL, cuBLAS and SYCL
by: Torres, L. A., et al.
Published: (2024)
by: Torres, L. A., et al.
Published: (2024)
RandNet-Parareal: a time-parallel PDE solver using Random Neural Networks
by: Gattiglio, Guglielmo, et al.
Published: (2024)
by: Gattiglio, Guglielmo, et al.
Published: (2024)
Massively Parallel Genetic Optimization through Asynchronous Propagation of Populations
by: Taubert, Oskar, et al.
Published: (2023)
by: Taubert, Oskar, et al.
Published: (2023)
Simopt-Power: Leveraging Simulation Metadata for Low-Power Design Synthesis
by: Wadhwa, Eashan, et al.
Published: (2025)
by: Wadhwa, Eashan, et al.
Published: (2025)
DP-HLS: A High-Level Synthesis Framework for Accelerating Dynamic Programming Algorithms in Bioinformatics
by: Cao, Yingqi, et al.
Published: (2024)
by: Cao, Yingqi, et al.
Published: (2024)
Global Optimizations & Lightweight Dynamic Logic for Concurrency
by: Pati, Suchita, et al.
Published: (2024)
by: Pati, Suchita, et al.
Published: (2024)
Parallel performance of shared memory parallel spectral deferred corrections
by: Freese, Philip, et al.
Published: (2024)
by: Freese, Philip, et al.
Published: (2024)
CPU-less parallel execution of lambda calculus in digital logic
by: Fitchett, Harry, et al.
Published: (2026)
by: Fitchett, Harry, et al.
Published: (2026)
Cascaded Prediction and Asynchronous Execution of Iterative Algorithms on Heterogeneous Platforms
by: Gao, Jianhua, et al.
Published: (2024)
by: Gao, Jianhua, et al.
Published: (2024)
Scalable Multilevel Monte Carlo Methods Exploiting Parallel Redistribution on Coarse Levels
by: Fairbanks, Hillary R., et al.
Published: (2024)
by: Fairbanks, Hillary R., et al.
Published: (2024)
The Detection and Correction of Silent Errors in Pipelined Krylov Subspace Methods
by: Carson, Erin Claire, et al.
Published: (2024)
by: Carson, Erin Claire, et al.
Published: (2024)
Fast and forward stable randomized algorithms for linear least-squares problems
by: Epperly, Ethan N.
Published: (2023)
by: Epperly, Ethan N.
Published: (2023)
Compute Can't Handle the Truth: Why Communication Tax Prioritizes Memory and Interconnects in Modern AI Infrastructure
by: Jung, Myoungsoo
Published: (2025)
by: Jung, Myoungsoo
Published: (2025)
GPU-Augmented OLAP Execution Engine: GPU Offloading
by: Chang, Ilsun
Published: (2025)
by: Chang, Ilsun
Published: (2025)
TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI
by: Kwak, Hyunseok, et al.
Published: (2025)
by: Kwak, Hyunseok, et al.
Published: (2025)
Similar Items
-
Regular mixed-radix DFT matrix factorization for in-place FFT accelerators
by: Salishev, Sergey
Published: (2025) -
PackSELL: A Sparse Matrix Format for Precision-Agnostic High-Performance SpMV
by: Suzuki, Kengo, et al.
Published: (2026) -
Wattchmen: Watching the Wattchers -- High Fidelity, Flexible GPU Energy Modeling
by: Tran, Brandon, et al.
Published: (2026) -
Exploiting nested task-parallelism in the $\mathcal{H}-LU$ factorization
by: Carratalá-Sáez, Rocío, et al.
Published: (2019) -
NCCL EP: Towards a Unified Expert Parallel Communication API for NCCL
by: Goldman, Amos, et al.
Published: (2026)