Cascaded Prediction and Asynchronous Execution of Iterative Algorithms on Heterogeneous Platforms
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Jianhua, Liu, Bingjie, Wang, Yizhuo, Ji, Weixing, Huang, Hua |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Systematic Literature Survey of Sparse Matrix-Vector Multiplication
by: Gao, Jianhua, et al.
Published: (2024)
by: Gao, Jianhua, et al.
Published: (2024)
Precision-Aware Iterative Algorithms Based on Group-Shared Exponents of Floating-Point Numbers
by: Gao, Jianhua, et al.
Published: (2024)
by: Gao, Jianhua, et al.
Published: (2024)
Efficient Parallel Scheduling for Sparse Triangular Solvers
by: Böhnlein, Toni, et al.
Published: (2025)
by: Böhnlein, Toni, et al.
Published: (2025)
Dynamic Memory Management on GPUs with SYCL
by: Standish, Russell K.
Published: (2025)
by: Standish, Russell K.
Published: (2025)
PackSELL: A Sparse Matrix Format for Precision-Agnostic High-Performance SpMV
by: Suzuki, Kengo, et al.
Published: (2026)
by: Suzuki, Kengo, et al.
Published: (2026)
Communication-Efficient, 2D Parallel Stochastic Gradient Descent for Distributed-Memory Optimization
by: Devarakonda, Aditya, et al.
Published: (2025)
by: Devarakonda, Aditya, et al.
Published: (2025)
Design, Configuration, Implementation, and Performance of a Simple 32 Core Raspberry Pi Cluster
by: Cicirello, Vincent A.
Published: (2017)
by: Cicirello, Vincent A.
Published: (2017)
Racing to Idle: Energy Efficiency of Matrix Multiplication on Heterogeneous CPU and GPU Architectures
by: Ansari, Mufakir Qamar, et al.
Published: (2025)
by: Ansari, Mufakir Qamar, et al.
Published: (2025)
Solving Large Rank-Deficient Linear Least-Squares Problems on Shared-Memory CPU Architectures and GPU Architectures
by: Chillarón, Mónica, et al.
Published: (2024)
by: Chillarón, Mónica, et al.
Published: (2024)
Accelerating Matrix Multiplication: A Performance Comparison Between Multi-Core CPU and GPU
by: Ansari, Mufakir Qamar, et al.
Published: (2025)
by: Ansari, Mufakir Qamar, et al.
Published: (2025)
Scalable Domain-decomposed Monte Carlo Neutral Transport for Nuclear Fusion
by: Lappi, Oskar, et al.
Published: (2025)
by: Lappi, Oskar, et al.
Published: (2025)
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation
by: Cheng, Long, et al.
Published: (2026)
by: Cheng, Long, et al.
Published: (2026)
Understanding GEMM Performance and Energy on NVIDIA Ada Lovelace: A Machine Learning-Based Analytical Approach
by: Xiaoteng, et al.
Published: (2024)
by: Xiaoteng, et al.
Published: (2024)
Stochastic well-structured transition systems
by: Aspnes, James
Published: (2025)
by: Aspnes, James
Published: (2025)
Utilizing Sparsity in the GPU-accelerated Assembly of Schur Complement Matrices in Domain Decomposition Methods
by: Homola, Jakub, et al.
Published: (2025)
by: Homola, Jakub, et al.
Published: (2025)
Gradient Coding with Iterative Block Leverage Score Sampling
by: Charalambides, Neophytos, et al.
Published: (2023)
by: Charalambides, Neophytos, et al.
Published: (2023)
Minimum Cost Loop Nests for Contraction of a Sparse Tensor with a Tensor Network
by: Kanakagiri, Raghavendra, et al.
Published: (2023)
by: Kanakagiri, Raghavendra, et al.
Published: (2023)
A Lock-Free, Fully GPU-Resident Architecture for the Verification of Goldbach's Conjecture
by: Llorente-Saguer, Isaac
Published: (2026)
by: Llorente-Saguer, Isaac
Published: (2026)
Floating Point Compression of Hierarchical Matrix Formats and its Impact on Matrix-Vector Multiplication
by: Kriemann, Ronald
Published: (2024)
by: Kriemann, Ronald
Published: (2024)
Distributed Hybrid Sketching for $\ell_2$-Embeddings
by: Charalambides, Neophytos, et al.
Published: (2024)
by: Charalambides, Neophytos, et al.
Published: (2024)
Stream parallel skeleton optimization
by: Aldinucci, Marco, et al.
Published: (2024)
by: Aldinucci, Marco, et al.
Published: (2024)
StreamFlow: cross-breeding cloud with HPC
by: Colonnelli, Iacopo, et al.
Published: (2020)
by: Colonnelli, Iacopo, et al.
Published: (2020)
A Study of Performance Portability in Plasma Physics Simulations
by: Ruzicka, Josef, et al.
Published: (2024)
by: Ruzicka, Josef, et al.
Published: (2024)
Optimizing Fine-Grained Parallelism Through Dynamic Load Balancing on Multi-Socket Many-Core Systems
by: Wang, Wenyi, et al.
Published: (2025)
by: Wang, Wenyi, et al.
Published: (2025)
Joint Training on AMD and NVIDIA GPUs
by: Hu, Jon, et al.
Published: (2026)
by: Hu, Jon, et al.
Published: (2026)
AutoTSMM: An Auto-tuning Framework for Building High-Performance Tall-and-Skinny Matrix-Matrix Multiplication on CPUs
by: Li, Chendi, et al.
Published: (2022)
by: Li, Chendi, et al.
Published: (2022)
Enabling Practical Transparent Checkpointing for MPI: A Topological Sort Approach
by: Xu, Yao, et al.
Published: (2024)
by: Xu, Yao, et al.
Published: (2024)
Leveraging Multi-Instance GPUs through moldable task scheduling
by: Villarrubia, Jorge, et al.
Published: (2025)
by: Villarrubia, Jorge, et al.
Published: (2025)
Data Scheduling Algorithm for Scalable and Efficient IoT Sensing in Cloud Computing
by: Mohammad, Noor Islam S.
Published: (2025)
by: Mohammad, Noor Islam S.
Published: (2025)
FlashSpread: IO-Aware GPU Simulation of Non-Markovian Epidemic Dynamics via Kernel Fusion
by: Shakeri, Heman, et al.
Published: (2026)
by: Shakeri, Heman, et al.
Published: (2026)
Declarative distributed algorithms as axiomatic theories in three-valued modal logic over semitopologies
by: Gabbay, Murdoch J.
Published: (2025)
by: Gabbay, Murdoch J.
Published: (2025)
Massively Parallel Genetic Optimization through Asynchronous Propagation of Populations
by: Taubert, Oskar, et al.
Published: (2023)
by: Taubert, Oskar, et al.
Published: (2023)
Canonicalization of Batched Einstein Summations for Tuning Retrieval
by: Kulkarni, Kaushik, et al.
Published: (2026)
by: Kulkarni, Kaushik, et al.
Published: (2026)
Challenging Portability Paradigms: FPGA Acceleration Using SYCL and OpenCL
by: de Castro, Manuel, et al.
Published: (2024)
by: de Castro, Manuel, et al.
Published: (2024)
A Hybrid Direct-Iterative Method for Solving KKT Linear Systems
by: Regev, Shaked, et al.
Published: (2021)
by: Regev, Shaked, et al.
Published: (2021)
A C++17 Thread Pool for High-Performance Scientific Computing
by: Shoshany, Barak
Published: (2021)
by: Shoshany, Barak
Published: (2021)
Population Protocols Revisited: Parity and Beyond
by: Gąsieniec, Leszek, et al.
Published: (2025)
by: Gąsieniec, Leszek, et al.
Published: (2025)
Distortion Resilience for Goal-Oriented Semantic Communication
by: Nguyen, Minh-Duong, et al.
Published: (2023)
by: Nguyen, Minh-Duong, et al.
Published: (2023)
DNA sequence alignment: An assignment for OpenMP, MPI, and CUDA/OpenCL
by: Gonzalez-Escribano, Arturo, et al.
Published: (2024)
by: Gonzalez-Escribano, Arturo, et al.
Published: (2024)
Distributed Tomographic Reconstruction with Quantization
by: Miao, Runxuan, et al.
Published: (2024)
by: Miao, Runxuan, et al.
Published: (2024)
Similar Items
-
A Systematic Literature Survey of Sparse Matrix-Vector Multiplication
by: Gao, Jianhua, et al.
Published: (2024) -
Precision-Aware Iterative Algorithms Based on Group-Shared Exponents of Floating-Point Numbers
by: Gao, Jianhua, et al.
Published: (2024) -
Efficient Parallel Scheduling for Sparse Triangular Solvers
by: Böhnlein, Toni, et al.
Published: (2025) -
Dynamic Memory Management on GPUs with SYCL
by: Standish, Russell K.
Published: (2025) -
PackSELL: A Sparse Matrix Format for Precision-Agnostic High-Performance SpMV
by: Suzuki, Kengo, et al.
Published: (2026)