Saved in:
| Main Authors: | Carrica, Vicki, Alomairy, Rabab, Ringoot, Evelyne, Edelman, Alan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2601.08082 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Toward Portable GPU Performance: Julia Recursive Implementation of TRMM and TRSM
by: Carrica, Vicki, et al.
Published: (2025)
by: Carrica, Vicki, et al.
Published: (2025)
Accelerating Bidiagonalization of Banded Matrices through Memory-Aware Bulge-Chasing on GPUs
by: Ringoot, Evelyne, et al.
Published: (2025)
by: Ringoot, Evelyne, et al.
Published: (2025)
Performant Unified GPU Kernels for Portable Singular Value Computation Across Hardware and Precision
by: Ringoot, Evelyne, et al.
Published: (2025)
by: Ringoot, Evelyne, et al.
Published: (2025)
Evaluating Fault Tolerance and Scalability in Distributed File Systems: A Case Study of GFS, HDFS, and MinIO
by: Malhotra, Shubham, et al.
Published: (2025)
by: Malhotra, Shubham, et al.
Published: (2025)
Optimizing Intra-Container Communication with Memory Protection Keys: A Novel Approach to Secure and Efficient Microservice Interaction
by: Yashu, Fnu, et al.
Published: (2025)
by: Yashu, Fnu, et al.
Published: (2025)
Seamless acceleration of Fortran intrinsics via AMD AI engines
by: Brown, Nick, et al.
Published: (2025)
by: Brown, Nick, et al.
Published: (2025)
The Landscape of GPU-Centric Communication
by: Unat, Didem, et al.
Published: (2024)
by: Unat, Didem, et al.
Published: (2024)
Active Inference-Based Adaptive Routing for Heterogeneous Edge AI Services
by: Wang, Zihang, et al.
Published: (2026)
by: Wang, Zihang, et al.
Published: (2026)
On the Performance of Cloud-based ARM SVE for Zero-Knowledge Proving Systems
by: Loghin, Dumitrel, et al.
Published: (2025)
by: Loghin, Dumitrel, et al.
Published: (2025)
A Communication Avoiding and Reducing Algorithm for Symmetric Eigenproblem for Very Small Matrices
by: Katagiri, Takahiro, et al.
Published: (2024)
by: Katagiri, Takahiro, et al.
Published: (2024)
A Hybrid Heuristic Framework for Resource-Efficient Querying of Scientific Experiments Data
by: Patel, Mayank, et al.
Published: (2025)
by: Patel, Mayank, et al.
Published: (2025)
Accelerating High-Order Finite Element Simulations at Extreme Scale with FP64 Tensor Cores
by: Tu, Jiqun, et al.
Published: (2026)
by: Tu, Jiqun, et al.
Published: (2026)
The Sunk Carbon Fallacy: Rethinking Carbon Footprint Metrics for Effective Carbon-Aware Scheduling
by: Bashir, Noman, et al.
Published: (2024)
by: Bashir, Noman, et al.
Published: (2024)
KPI2KVI: A Multi Agent Workflow for Calculating Key Value Indicators from Service Descriptions
by: Shokrnezhad, Masoud, et al.
Published: (2026)
by: Shokrnezhad, Masoud, et al.
Published: (2026)
Wattlytics: A Web Platform for Co-Optimizing Performance, Energy, and TCO in HPC Clusters
by: Afzal, Ayesha, et al.
Published: (2026)
by: Afzal, Ayesha, et al.
Published: (2026)
E-QUARTIC: Energy Efficient Edge Ensemble of Convolutional Neural Networks for Resource-Optimized Learning
by: Zhang, Le, et al.
Published: (2024)
by: Zhang, Le, et al.
Published: (2024)
pyGinkgo: A Sparse Linear Algebra Operator Framework for Python
by: Tuteja, Keshvi, et al.
Published: (2025)
by: Tuteja, Keshvi, et al.
Published: (2025)
Model Discovery and Graph Simulation: A Lightweight Gateway to Chaos Engineering
by: Krasnovsky, Anatoly A.
Published: (2025)
by: Krasnovsky, Anatoly A.
Published: (2025)
Xabclib:A Fully Auto-tuned Sparse Iterative Solver
by: Katagiri, Takahiro, et al.
Published: (2024)
by: Katagiri, Takahiro, et al.
Published: (2024)
Automated MPI-X code generation for scalable finite-difference solvers
by: Bisbas, George, et al.
Published: (2023)
by: Bisbas, George, et al.
Published: (2023)
On the energy efficiency of sparse matrix computations on multi-GPU clusters
by: Bernaschi, Massimo, et al.
Published: (2025)
by: Bernaschi, Massimo, et al.
Published: (2025)
NApy: Efficient Statistics in Python for Large-Scale Heterogeneous Data with Enhanced Support for Missing Data
by: Woller, Fabian, et al.
Published: (2025)
by: Woller, Fabian, et al.
Published: (2025)
Beating vDSP: A 138 GFLOPS Radix-8 Stockham FFT on Apple Silicon via Two-Tier Register-Threadgroup Memory Decomposition
by: Bergach, Mohamed Amine
Published: (2026)
by: Bergach, Mohamed Amine
Published: (2026)
Performance measurements of modern Fortran MPI applications with Score-P
by: Corbin, Gregor
Published: (2025)
by: Corbin, Gregor
Published: (2025)
Black-Scholes Option Pricing on Intel CPUs and GPUs: Implementation on SYCL and Optimization Techniques
by: Panova, Elena, et al.
Published: (2022)
by: Panova, Elena, et al.
Published: (2022)
GoldbachGPU: An Open Source GPU-Accelerated Framework for Verification of Goldbach's Conjecture
by: Llorente-Saguer, Isaac
Published: (2026)
by: Llorente-Saguer, Isaac
Published: (2026)
Performant Automatic BLAS Offloading on Unified Memory Architecture with OpenMP First-Touch Style Data Movement
by: Li, Junjie
Published: (2024)
by: Li, Junjie
Published: (2024)
Harnessing the Full Potential of RRAMs through Scalable and Distributed In-Memory Computing with Integrated Error Correction
by: Vo, Huynh Q. N., et al.
Published: (2025)
by: Vo, Huynh Q. N., et al.
Published: (2025)
Deep Optimizer States: Towards Scalable Training of Transformer Models Using Interleaved Offloading
by: Maurya, Avinash, et al.
Published: (2024)
by: Maurya, Avinash, et al.
Published: (2024)
Scaling Intelligence: Designing Data Centers for Next-Gen Language Models
by: Tithi, Jesmin Jahan, et al.
Published: (2025)
by: Tithi, Jesmin Jahan, et al.
Published: (2025)
SCALE: Self-regulated Clustered federAted LEarning in a Homogeneous Environment
by: Puppala, Sai, et al.
Published: (2024)
by: Puppala, Sai, et al.
Published: (2024)
Easy Acceleration with Distributed Arrays
by: Kepner, Jeremy, et al.
Published: (2025)
by: Kepner, Jeremy, et al.
Published: (2025)
Quantum resources in resource management systems
by: Bacher, Utz, et al.
Published: (2025)
by: Bacher, Utz, et al.
Published: (2025)
Execution Envelopes: A Shared Admission Contract for Backend AI Execution Requests
by: Tallam, Krti
Published: (2026)
by: Tallam, Krti
Published: (2026)
Optimizing Spot Instance Reliability and Security Using Cloud-Native Data and Tools
by: Saqib, Muhammad, et al.
Published: (2025)
by: Saqib, Muhammad, et al.
Published: (2025)
Generative AI for Software Architecture. Applications, Challenges, and Future Directions
by: Esposito, Matteo, et al.
Published: (2025)
by: Esposito, Matteo, et al.
Published: (2025)
Quantum-HPC Software Stacks and the openQSE Reference Architecture: A Survey
by: Shehata, Amir, et al.
Published: (2026)
by: Shehata, Amir, et al.
Published: (2026)
Scaling Sample-Based Quantum Diagonalization on GPU-Accelerated Systems using OpenMP Offload
by: Walkup, Robert, et al.
Published: (2026)
by: Walkup, Robert, et al.
Published: (2026)
GPU-Accelerated Quantum Simulation: Empirical Backend Selection, Gate Fusion, and Adaptive Precision
by: Kumaresan, Poornima, et al.
Published: (2026)
by: Kumaresan, Poornima, et al.
Published: (2026)
From GPUs to RRAMs: Distributed In-Memory Primal-Dual Hybrid Gradient Method for Solving Large-Scale Linear Optimization Problem
by: Vo, Huynh Q. N., et al.
Published: (2025)
by: Vo, Huynh Q. N., et al.
Published: (2025)
Similar Items
-
Toward Portable GPU Performance: Julia Recursive Implementation of TRMM and TRSM
by: Carrica, Vicki, et al.
Published: (2025) -
Accelerating Bidiagonalization of Banded Matrices through Memory-Aware Bulge-Chasing on GPUs
by: Ringoot, Evelyne, et al.
Published: (2025) -
Performant Unified GPU Kernels for Portable Singular Value Computation Across Hardware and Precision
by: Ringoot, Evelyne, et al.
Published: (2025) -
Evaluating Fault Tolerance and Scalability in Distributed File Systems: A Case Study of GFS, HDFS, and MinIO
by: Malhotra, Shubham, et al.
Published: (2025) -
Optimizing Intra-Container Communication with Memory Protection Keys: A Novel Approach to Secure and Efficient Microservice Interaction
by: Yashu, Fnu, et al.
Published: (2025)