JAXMg: A multi-GPU linear solver in JAX
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Wiersema, Roeland |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DrJAX: Scalable and Differentiable MapReduce Primitives in JAX
von: Rush, Keith, et al.
Veröffentlicht: (2024)
von: Rush, Keith, et al.
Veröffentlicht: (2024)
Orbax: Distributed Checkpointing with JAX
von: Gaffney, Colin, et al.
Veröffentlicht: (2026)
von: Gaffney, Colin, et al.
Veröffentlicht: (2026)
Profiling and optimization of multi-card GPU machine learning jobs
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
Testing and benchmarking emerging supercomputers via the MFC flow solver
von: Wilfong, Benjamin, et al.
Veröffentlicht: (2025)
von: Wilfong, Benjamin, et al.
Veröffentlicht: (2025)
Improving GPU Multi-Tenancy Through Dynamic Multi-Instance GPU Reconfiguration
von: Wang, Tianyu, et al.
Veröffentlicht: (2024)
von: Wang, Tianyu, et al.
Veröffentlicht: (2024)
Efficient allocation of image recognition and LLM tasks on multi-GPU system
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
Regent based parallel meshfree LSKUM solver for heterogenous HPC platforms
von: Salil, Sanath, et al.
Veröffentlicht: (2024)
von: Salil, Sanath, et al.
Veröffentlicht: (2024)
Performance of a high-order MPI-Kokkos accelerated fluid solver
von: Sporykhin, Filipp, et al.
Veröffentlicht: (2025)
von: Sporykhin, Filipp, et al.
Veröffentlicht: (2025)
Workflow decomposition algorithm for scheduling with quantum annealer-based hybrid solver
von: Kroczek, Marcin, et al.
Veröffentlicht: (2025)
von: Kroczek, Marcin, et al.
Veröffentlicht: (2025)
Accelerating Biclique Counting on GPU
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
GPU Sharing with Triples Mode
von: Byun, Chansup, et al.
Veröffentlicht: (2024)
von: Byun, Chansup, et al.
Veröffentlicht: (2024)
Accelerating Intra-Node GPU-to-GPU Communication Through Multi-Path Transfers with CUDA Graphs
von: Sojoodi, Amirhossein, et al.
Veröffentlicht: (2026)
von: Sojoodi, Amirhossein, et al.
Veröffentlicht: (2026)
ParvaGPU: Efficient Spatial GPU Sharing for Large-Scale DNN Inference in Cloud Environments
von: Lee, Munkyu, et al.
Veröffentlicht: (2024)
von: Lee, Munkyu, et al.
Veröffentlicht: (2024)
A Multi-Objective Framework for Optimizing GPU-Enabled VM Placement in Cloud Data Centers with Multi-Instance GPU Technology
von: Siavashi, Ahmad, et al.
Veröffentlicht: (2025)
von: Siavashi, Ahmad, et al.
Veröffentlicht: (2025)
Breaking the Memory Wall: A Study of I/O Patterns and GPU Memory Utilization for Hybrid CPU-GPU Offloaded Optimizers
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
A Unified CPU-GPU Protocol for GNN Training
von: Lin, Yi-Chien, et al.
Veröffentlicht: (2024)
von: Lin, Yi-Chien, et al.
Veröffentlicht: (2024)
DuaLip-GPU Technical Report
von: Dexter, Gregory, et al.
Veröffentlicht: (2026)
von: Dexter, Gregory, et al.
Veröffentlicht: (2026)
Incidence Constraints in Hypergraph Partitioning on GPU
von: Ronzani, Marco, et al.
Veröffentlicht: (2026)
von: Ronzani, Marco, et al.
Veröffentlicht: (2026)
Predictable LLM Serving on GPU Clusters
von: Darzi, Erfan, et al.
Veröffentlicht: (2025)
von: Darzi, Erfan, et al.
Veröffentlicht: (2025)
GPU Accelerated Sparse Cholesky Factorization
von: Karsavuran, M. Ozan, et al.
Veröffentlicht: (2024)
von: Karsavuran, M. Ozan, et al.
Veröffentlicht: (2024)
Heat: Satellite's meat is GPU's poison
von: Yuan, Zhehu, et al.
Veröffentlicht: (2024)
von: Yuan, Zhehu, et al.
Veröffentlicht: (2024)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
von: Gu, Jianfeng, et al.
Veröffentlicht: (2025)
von: Gu, Jianfeng, et al.
Veröffentlicht: (2025)
A Preliminary Study on Accelerating Simulation Optimization with GPU Implementation
von: He, Jinghai, et al.
Veröffentlicht: (2024)
von: He, Jinghai, et al.
Veröffentlicht: (2024)
A Framework for Fine-Grained Synchronization of Dependent GPU Kernels
von: Jangda, Abhinav, et al.
Veröffentlicht: (2023)
von: Jangda, Abhinav, et al.
Veröffentlicht: (2023)
TURNIP: A "Nondeterministic" GPU Runtime with CPU RAM Offload
von: Ding, Zhimin, et al.
Veröffentlicht: (2024)
von: Ding, Zhimin, et al.
Veröffentlicht: (2024)
A Practical GPU-Accelerated Implementation of Orthogonal Matching Pursuit
von: Lubonja, Ariel, et al.
Veröffentlicht: (2024)
von: Lubonja, Ariel, et al.
Veröffentlicht: (2024)
Optimizing Bloom Filters for Modern GPU Architectures
von: Jünger, Daniel, et al.
Veröffentlicht: (2025)
von: Jünger, Daniel, et al.
Veröffentlicht: (2025)
CRIUgpu: Transparent Checkpointing of GPU-Accelerated Workloads
von: Stoyanov, Radostin, et al.
Veröffentlicht: (2025)
von: Stoyanov, Radostin, et al.
Veröffentlicht: (2025)
Understanding the Landscape of Ampere GPU Memory Errors
von: Zhu, Zhu, et al.
Veröffentlicht: (2025)
von: Zhu, Zhu, et al.
Veröffentlicht: (2025)
Methodology for GPU Frequency Switching Latency Measurement
von: Velicka, Daniel, et al.
Veröffentlicht: (2025)
von: Velicka, Daniel, et al.
Veröffentlicht: (2025)
Adaptive Multidimensional Quadrature on Multi-GPU Systems
von: Tonarelli, Melanie, et al.
Veröffentlicht: (2025)
von: Tonarelli, Melanie, et al.
Veröffentlicht: (2025)
GPU-Accelerated Batch-Dynamic Subgraph Matching
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
MERBIT: A GPU-Based SpMV Method for Iterative Workloads
von: Zhang, Qi, et al.
Veröffentlicht: (2026)
von: Zhang, Qi, et al.
Veröffentlicht: (2026)
A GPU Accelerated Temporal Window-Based Random Walk Sampler
von: Salehin, Md Ashfaq, et al.
Veröffentlicht: (2026)
von: Salehin, Md Ashfaq, et al.
Veröffentlicht: (2026)
A GPU-accelerated Molecular Docking Workflow with Kubernetes and Apache Airflow
von: Medeiros, Daniel, et al.
Veröffentlicht: (2024)
von: Medeiros, Daniel, et al.
Veröffentlicht: (2024)
Characterization-Guided GPU Fault Resilience in NVIDIA MPS
von: Liu, Rixin, et al.
Veröffentlicht: (2026)
von: Liu, Rixin, et al.
Veröffentlicht: (2026)
Hummingbird: SLO-Oriented GPU Preemption at Microsecond-scale
von: Hu, Tiancheng, et al.
Veröffentlicht: (2026)
von: Hu, Tiancheng, et al.
Veröffentlicht: (2026)
Efficient Accelerated Graph Edit Distance Computation on GPU
von: Dabah, Adel, et al.
Veröffentlicht: (2026)
von: Dabah, Adel, et al.
Veröffentlicht: (2026)
VDCores: Resource Decoupled Programming and Execution for Asynchronous GPU
von: He, Zijian, et al.
Veröffentlicht: (2026)
von: He, Zijian, et al.
Veröffentlicht: (2026)
Accelerating Sparse MTTKRP for Small Tensor Decomposition on GPU
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DrJAX: Scalable and Differentiable MapReduce Primitives in JAX
von: Rush, Keith, et al.
Veröffentlicht: (2024) -
Orbax: Distributed Checkpointing with JAX
von: Gaffney, Colin, et al.
Veröffentlicht: (2026) -
Profiling and optimization of multi-card GPU machine learning jobs
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025) -
Testing and benchmarking emerging supercomputers via the MFC flow solver
von: Wilfong, Benjamin, et al.
Veröffentlicht: (2025) -
Improving GPU Multi-Tenancy Through Dynamic Multi-Instance GPU Reconfiguration
von: Wang, Tianyu, et al.
Veröffentlicht: (2024)