HYLU: Hybrid Parallel Sparse LU Factorization
Fuente:
arXiv
Saved in:
| Main Author: | Chen, Xiaoming |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Utilizing Sparsity in the GPU-accelerated Assembly of Schur Complement Matrices in Domain Decomposition Methods
by: Homola, Jakub, et al.
Published: (2025)
by: Homola, Jakub, et al.
Published: (2025)
On Advanced Monte Carlo Methods for Linear Algebra on Advanced Accelerator Architectures
by: Lebedev, Anton, et al.
Published: (2024)
by: Lebedev, Anton, et al.
Published: (2024)
Parallel Sparse and Data-Sparse Factorization-based Linear Solvers
by: Li, Xiaoye Sherry, et al.
Published: (2026)
by: Li, Xiaoye Sherry, et al.
Published: (2026)
Evaluation of POSIT Arithmetic with Accelerators
by: Nakasato, Naohito, et al.
Published: (2024)
by: Nakasato, Naohito, et al.
Published: (2024)
Fast and energy-efficient derivatives risk analysis: Streaming option Greeks on Xilinx and Intel FPGAs
by: Klaisoongnoen, Mark, et al.
Published: (2022)
by: Klaisoongnoen, Mark, et al.
Published: (2022)
Minimum Cost Loop Nests for Contraction of a Sparse Tensor with a Tensor Network
by: Kanakagiri, Raghavendra, et al.
Published: (2023)
by: Kanakagiri, Raghavendra, et al.
Published: (2023)
Canonicalization of Batched Einstein Summations for Tuning Retrieval
by: Kulkarni, Kaushik, et al.
Published: (2026)
by: Kulkarni, Kaushik, et al.
Published: (2026)
Distributed-memory Algorithms for Sparse Matrix Permutation, Extraction, and Assignment
by: Hassani, Elaheh, et al.
Published: (2025)
by: Hassani, Elaheh, et al.
Published: (2025)
TriADA: Massively Parallel Trilinear Matrix-by-Tensor Multiply-Add Algorithm and Device Architecture for the Acceleration of 3D Discrete Transformations
by: Sedukhin, Stanislav, et al.
Published: (2025)
by: Sedukhin, Stanislav, et al.
Published: (2025)
pLUTo: Enabling Massively Parallel Computation in DRAM via Lookup Tables
by: Ferreira, João Dinis, et al.
Published: (2021)
by: Ferreira, João Dinis, et al.
Published: (2021)
Efficient Hardware Accelerator Based on Medium Granularity Dataflow for SpTRSV
by: Chen, Qian, et al.
Published: (2024)
by: Chen, Qian, et al.
Published: (2024)
Datapath Combinational Equivalence Checking With Hybrid Sweeping Engines and Parallelization
by: Chen, Zhihan, et al.
Published: (2024)
by: Chen, Zhihan, et al.
Published: (2024)
Using matrices in post-processing phase of CFD simulations
by: Argentini, Gianluca
Published: (2004)
by: Argentini, Gianluca
Published: (2004)
A comprehensive evaluation of spatial co-execution on GPUs using MPS and MIG technologies
by: Villarrubia, Jorge, et al.
Published: (2026)
by: Villarrubia, Jorge, et al.
Published: (2026)
Tensor Decompositions for Count Data that Leverage Stochastic and Deterministic Optimization
by: Myers, Jeremy M., et al.
Published: (2022)
by: Myers, Jeremy M., et al.
Published: (2022)
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation
by: Cheng, Long, et al.
Published: (2026)
by: Cheng, Long, et al.
Published: (2026)
Speed, power and cost implications for GPU acceleration of Computational Fluid Dynamics on HPC systems
by: Cooper-Baldock, Zachary, et al.
Published: (2024)
by: Cooper-Baldock, Zachary, et al.
Published: (2024)
HieraSparse: Hierarchical Semi-Structured Sparse KV Attention
by: Wang, Haoxuan, et al.
Published: (2026)
by: Wang, Haoxuan, et al.
Published: (2026)
TREA: Low-precision Time-Multiplexed, Resource-Efficient Edge Accelerator for Object Detection and Classification
by: Sharma, Vijay Pratap, et al.
Published: (2026)
by: Sharma, Vijay Pratap, et al.
Published: (2026)
Regular mixed-radix DFT matrix factorization for in-place FFT accelerators
by: Salishev, Sergey
Published: (2025)
by: Salishev, Sergey
Published: (2025)
Communication-Efficient and Memory-Aware Parallel Bootstrapping using MPI
by: Zhang, Di
Published: (2025)
by: Zhang, Di
Published: (2025)
Parendi: Thousand-Way Parallel RTL Simulation
by: Emami, Mahyar, et al.
Published: (2024)
by: Emami, Mahyar, et al.
Published: (2024)
Exploring the Design Space for Message-Driven Systems for Dynamic Graph Processing using CCA
by: Chandio, Bibrak Qamar, et al.
Published: (2024)
by: Chandio, Bibrak Qamar, et al.
Published: (2024)
Deep Learning and Machine Learning with GPGPU and CUDA: Unlocking the Power of Parallel Computing
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
by: Zhang, Chen, et al.
Published: (2026)
by: Zhang, Chen, et al.
Published: (2026)
Sparse MTTKRP Acceleration for Tensor Decomposition on GPU
by: Wijeratne, Sasindu, et al.
Published: (2024)
by: Wijeratne, Sasindu, et al.
Published: (2024)
SpArch: Efficient Architecture for Sparse Matrix Multiplication
by: Zhang, Zhekai, et al.
Published: (2020)
by: Zhang, Zhekai, et al.
Published: (2020)
How Fast Can Graph Computations Go on Fine-grained Parallel Architectures
by: Wang, Yuqing, et al.
Published: (2025)
by: Wang, Yuqing, et al.
Published: (2025)
Decompose, Optimize, and Reconstruct: Very Large Constant Multiplication at Scale
by: Cantaloube, Théo, et al.
Published: (2026)
by: Cantaloube, Théo, et al.
Published: (2026)
Toward a Universal GPU Instruction Set Architecture: A Cross-Vendor Analysis of Hardware-Invariant Computational Primitives in Parallel Processors
by: Abraham, Ojima, et al.
Published: (2026)
by: Abraham, Ojima, et al.
Published: (2026)
Optimized thread-block arrangement in a GPU implementation of a linear solver for atmospheric chemistry mechanisms
by: Ruiz, Christian Guzman, et al.
Published: (2024)
by: Ruiz, Christian Guzman, et al.
Published: (2024)
Taming Offload Overheads in a Massively Parallel Open-Source RISC-V MPSoC: Analysis and Optimization
by: Colagrande, Luca, et al.
Published: (2025)
by: Colagrande, Luca, et al.
Published: (2025)
SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
by: Xu, Weihong, et al.
Published: (2025)
by: Xu, Weihong, et al.
Published: (2025)
Efficient Parallel Scheduling for Sparse Triangular Solvers
by: Böhnlein, Toni, et al.
Published: (2025)
by: Böhnlein, Toni, et al.
Published: (2025)
iHAC: A Hybrid Cluster Architecture for Enhanced Performance and Resilience
by: Muntaka, Siddique Abubakr, et al.
Published: (2026)
by: Muntaka, Siddique Abubakr, et al.
Published: (2026)
DFabric: Scaling Out Data Parallel Applications with CXL-Ethernet Hybrid Interconnects
by: Zhang, Xu, et al.
Published: (2024)
by: Zhang, Xu, et al.
Published: (2024)
DUET: Disaggregated Hybrid Mamba-Transformer LLMs with Prefill and Decode-Specific Packages
by: Kanani, Alish, et al.
Published: (2026)
by: Kanani, Alish, et al.
Published: (2026)
GigaAPI for GPU Parallelization
by: Suvarna, M., et al.
Published: (2025)
by: Suvarna, M., et al.
Published: (2025)
CELLO: Co-designing Schedule and Hybrid Implicit/Explicit Buffer for Complex Tensor Reuse
by: Garg, Raveesh, et al.
Published: (2023)
by: Garg, Raveesh, et al.
Published: (2023)
Lincoln AI Computing Survey (LAICS) and Trends
by: Reuther, Albert, et al.
Published: (2025)
by: Reuther, Albert, et al.
Published: (2025)
Similar Items
-
Utilizing Sparsity in the GPU-accelerated Assembly of Schur Complement Matrices in Domain Decomposition Methods
by: Homola, Jakub, et al.
Published: (2025) -
On Advanced Monte Carlo Methods for Linear Algebra on Advanced Accelerator Architectures
by: Lebedev, Anton, et al.
Published: (2024) -
Parallel Sparse and Data-Sparse Factorization-based Linear Solvers
by: Li, Xiaoye Sherry, et al.
Published: (2026) -
Evaluation of POSIT Arithmetic with Accelerators
by: Nakasato, Naohito, et al.
Published: (2024) -
Fast and energy-efficient derivatives risk analysis: Streaming option Greeks on Xilinx and Intel FPGAs
by: Klaisoongnoen, Mark, et al.
Published: (2022)