Pipelined Dense Symmetric Eigenvalue Decomposition on Multi-GPU Architectures
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Hansheng, Zhan, Ruiyi, Huang, Dajun, Liu, Xingchen, Li, Qiao, Duan, Hancong, Tao, Dingwen, Tan, Guangming, Zhang, Shaoshuai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Extracting the Potential of Emerging Hardware Accelerators for Symmetric Eigenvalue Decomposition
by: Wang, Hansheng, et al.
Published: (2024)
by: Wang, Hansheng, et al.
Published: (2024)
Implementing Multi-GPU Scientific Computing Miniapps Across Performance Portable Frameworks
by: Villalobos, Johansell, et al.
Published: (2025)
by: Villalobos, Johansell, et al.
Published: (2025)
Environmental Impact of CI/CD Pipelines
by: Saavedra, Nuno, et al.
Published: (2025)
by: Saavedra, Nuno, et al.
Published: (2025)
Towards an Optimized Benchmarking Platform for CI/CD Pipelines
by: Japke, Nils, et al.
Published: (2025)
by: Japke, Nils, et al.
Published: (2025)
Toward Portable GPU Performance: Julia Recursive Implementation of TRMM and TRSM
by: Carrica, Vicki, et al.
Published: (2025)
by: Carrica, Vicki, et al.
Published: (2025)
Ocean: Fast Estimation-Based Sparse General Matrix-Matrix Multiplication on GPU
by: Li, Yifan, et al.
Published: (2026)
by: Li, Yifan, et al.
Published: (2026)
Communication-Avoiding SpGEMM via Trident Partitioning on Hierarchical GPU Interconnects
by: Bellavita, Julian, et al.
Published: (2026)
by: Bellavita, Julian, et al.
Published: (2026)
Integrating Odeint Time Stepping into OpenFPM for Distributed and GPU Accelerated Numerical Solvers
by: Singh, Abhinav, et al.
Published: (2023)
by: Singh, Abhinav, et al.
Published: (2023)
Performant Unified GPU Kernels for Portable Singular Value Computation Across Hardware and Precision
by: Ringoot, Evelyne, et al.
Published: (2025)
by: Ringoot, Evelyne, et al.
Published: (2025)
Investigating Matrix Repartitioning to Address the Over- and Undersubscription Challenge for a GPU-based CFD Solver
by: Olenik, Gregor, et al.
Published: (2025)
by: Olenik, Gregor, et al.
Published: (2025)
TSUE: A Two-Stage Data Update Method for an Erasure Coded Cluster File System
by: Wei, Zheng, et al.
Published: (2025)
by: Wei, Zheng, et al.
Published: (2025)
Overcoming Memory Constraints in Quantum Circuit Simulation with a High-Fidelity Compression Framework
by: Zhang, Boyuan, et al.
Published: (2024)
by: Zhang, Boyuan, et al.
Published: (2024)
Robustness and Accuracy in Pipelined Bi-Conjugate Gradient Stabilized Method: A Comparative Study
by: Havdiak, Mykhailo, et al.
Published: (2024)
by: Havdiak, Mykhailo, et al.
Published: (2024)
ElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal Parallelism
by: Liu, Zedong, et al.
Published: (2025)
by: Liu, Zedong, et al.
Published: (2025)
An Analysis of HPC and Edge Architectures in the Cloud
by: Santillan, Steven, et al.
Published: (2025)
by: Santillan, Steven, et al.
Published: (2025)
CARISMA: CAR-Integrated Service Mesh Architecture
by: Klein, Kevin, et al.
Published: (2024)
by: Klein, Kevin, et al.
Published: (2024)
Proceedings First Workshop on Adaptable Cloud Architectures
by: De Palma, Giuseppe, et al.
Published: (2025)
by: De Palma, Giuseppe, et al.
Published: (2025)
A Reference Architecture for Governance of Cloud Native Applications
by: Pourmajidi, William, et al.
Published: (2023)
by: Pourmajidi, William, et al.
Published: (2023)
On the correlation between Architectural Smells and Static Analysis Warnings
by: Esposito, Matteo, et al.
Published: (2024)
by: Esposito, Matteo, et al.
Published: (2024)
A Scalable Clustered Architecture for Cyber-Physical Systems
by: Cabral, Bernardo
Published: (2024)
by: Cabral, Bernardo
Published: (2024)
Self-adaptive Multi-Access Edge Architectures: A Robotics Case
by: Moghaddam, Mahyar T, et al.
Published: (2026)
by: Moghaddam, Mahyar T, et al.
Published: (2026)
SoK: Microservice Architectures from a Dependability Perspective
by: Kažemaks, Dāvis, et al.
Published: (2025)
by: Kažemaks, Dāvis, et al.
Published: (2025)
TorchGWAS : GPU-accelerated GWAS for thousands of quantitative phenotypes
by: Zhao, Xingzhong, et al.
Published: (2026)
by: Zhao, Xingzhong, et al.
Published: (2026)
On the energy efficiency of sparse matrix computations on multi-GPU clusters
by: Bernaschi, Massimo, et al.
Published: (2025)
by: Bernaschi, Massimo, et al.
Published: (2025)
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
by: Nichols, Daniel, et al.
Published: (2025)
by: Nichols, Daniel, et al.
Published: (2025)
A Communication Avoiding and Reducing Algorithm for Symmetric Eigenproblem for Very Small Matrices
by: Katagiri, Takahiro, et al.
Published: (2024)
by: Katagiri, Takahiro, et al.
Published: (2024)
CSnake: Detecting Self-Sustaining Cascading Failure via Causal Stitching of Fault Propagations
by: Qian, Shangshu, et al.
Published: (2025)
by: Qian, Shangshu, et al.
Published: (2025)
A Framework for Effective Invocation Methods of Various LLM Services
by: Wang, Can, et al.
Published: (2024)
by: Wang, Can, et al.
Published: (2024)
GPU Implementations for Midsize Integer Addition and Multiplication
by: Oancea, Cosmin E., et al.
Published: (2024)
by: Oancea, Cosmin E., et al.
Published: (2024)
An LLVM-Based Optimization Pipeline for SPDZ
by: Dai, Tianye, et al.
Published: (2025)
by: Dai, Tianye, et al.
Published: (2025)
Multi-Grained Specifications for Distributed System Model Checking and Verification
by: Ouyang, Lingzhi, et al.
Published: (2024)
by: Ouyang, Lingzhi, et al.
Published: (2024)
Multi-Objective Load Balancing for Heterogeneous Edge-Based Object Detection Systems
by: Alqahtani, Daghash K., et al.
Published: (2026)
by: Alqahtani, Daghash K., et al.
Published: (2026)
SPUMA: a minimally invasive approach to the GPU porting of OPENFOAM
by: Bnà, Simone, et al.
Published: (2025)
by: Bnà, Simone, et al.
Published: (2025)
JanusPipe: Efficient Pipeline Parallel Training for Machine Learning Interatomic Potentials
by: Wang, Hongyu, et al.
Published: (2026)
by: Wang, Hongyu, et al.
Published: (2026)
Cost-Effective Big Data Orchestration Using Dagster: A Multi-Platform Approach
by: Picatto, Hernan, et al.
Published: (2024)
by: Picatto, Hernan, et al.
Published: (2024)
ContiguousKV: Accelerating LLM Prefill with Granularity-Aligned KV Cache Management
by: Zou, Jing, et al.
Published: (2026)
by: Zou, Jing, et al.
Published: (2026)
GoldbachGPU: An Open Source GPU-Accelerated Framework for Verification of Goldbach's Conjecture
by: Llorente-Saguer, Isaac
Published: (2026)
by: Llorente-Saguer, Isaac
Published: (2026)
Adapting Multi-objectivized Software Configuration Tuning
by: Chen, Tao, et al.
Published: (2024)
by: Chen, Tao, et al.
Published: (2024)
TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training
by: Liu, Man, et al.
Published: (2026)
by: Liu, Man, et al.
Published: (2026)
SGPRS: Seamless GPU Partitioning Real-Time Scheduler for Periodic Deep Learning Workloads
by: Babaei, Amir Fakhim, et al.
Published: (2024)
by: Babaei, Amir Fakhim, et al.
Published: (2024)
Similar Items
-
Extracting the Potential of Emerging Hardware Accelerators for Symmetric Eigenvalue Decomposition
by: Wang, Hansheng, et al.
Published: (2024) -
Implementing Multi-GPU Scientific Computing Miniapps Across Performance Portable Frameworks
by: Villalobos, Johansell, et al.
Published: (2025) -
Environmental Impact of CI/CD Pipelines
by: Saavedra, Nuno, et al.
Published: (2025) -
Towards an Optimized Benchmarking Platform for CI/CD Pipelines
by: Japke, Nils, et al.
Published: (2025) -
Toward Portable GPU Performance: Julia Recursive Implementation of TRMM and TRSM
by: Carrica, Vicki, et al.
Published: (2025)