A GPU accelerated mixed-precision Smoothed Particle Hydrodynamics framework with cell-based relative coordinates
Fuente:
arXiv
Saved in:
| Main Authors: | Mao, Zirui, Li, Xinyi, Hu, Shenyang, Gopalakrishnan, Ganesh, Li, Ang |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SIMT/GPU Data Race Verification using ISCC and Intermediary Code Representations: A Case Study
by: Osterhout, Andrew, et al.
Published: (2025)
by: Osterhout, Andrew, et al.
Published: (2025)
HiRace: Accurate and Fast Source-Level Race Checking of GPU Programs
by: Jacobson, John, et al.
Published: (2024)
by: Jacobson, John, et al.
Published: (2024)
GPU-Accelerated Distributed QAOA on Large-scale HPC Ecosystems
by: Xu, Zhihao, et al.
Published: (2025)
by: Xu, Zhihao, et al.
Published: (2025)
GPZ: GPU-Accelerated Lossy Compressor for Particle Data
by: Li, Ruoyu, et al.
Published: (2025)
by: Li, Ruoyu, et al.
Published: (2025)
The Shamrock code: I- Smoothed Particle Hydrodynamics on GPUs
by: David--Cléris, Timothée, et al.
Published: (2025)
by: David--Cléris, Timothée, et al.
Published: (2025)
Combining GPU and CPU for accelerating evolutionary computing workloads
by: Eynaliyev, Rustam, et al.
Published: (2025)
by: Eynaliyev, Rustam, et al.
Published: (2025)
A GPU-accelerated Molecular Docking Workflow with Kubernetes and Apache Airflow
by: Medeiros, Daniel, et al.
Published: (2024)
by: Medeiros, Daniel, et al.
Published: (2024)
Efficient GPU Implementation of Particle Interactions with Cutoff Radius and Few Particles per Cell
by: Algis, David, et al.
Published: (2024)
by: Algis, David, et al.
Published: (2024)
Characterization-Guided GPU Fault Resilience in NVIDIA MPS
by: Liu, Rixin, et al.
Published: (2026)
by: Liu, Rixin, et al.
Published: (2026)
A Study of Performance Programming of CPU, GPU accelerated Computers and SIMD Architecture
by: Yi, Xinyao
Published: (2024)
by: Yi, Xinyao
Published: (2024)
WgPy: GPU-accelerated NumPy-like array library for web browsers
by: Hidaka, Masatoshi, et al.
Published: (2025)
by: Hidaka, Masatoshi, et al.
Published: (2025)
Improving GPU Multi-Tenancy Through Dynamic Multi-Instance GPU Reconfiguration
by: Wang, Tianyu, et al.
Published: (2024)
by: Wang, Tianyu, et al.
Published: (2024)
Fast Topology-Aware Lossy Data Compression with Full Preservation of Critical Points and Local Order
by: Fallin, Alex, et al.
Published: (2026)
by: Fallin, Alex, et al.
Published: (2026)
TopoSZp: Lightweight Topology-Aware Error-controlled Compression for Scientific Data
by: Agarwal, Tripti, et al.
Published: (2026)
by: Agarwal, Tripti, et al.
Published: (2026)
Scalable GPU Performance Variability Analysis framework
by: Lahiry, Ankur, et al.
Published: (2025)
by: Lahiry, Ankur, et al.
Published: (2025)
FATE: Future-State-Aware Scheduling for Heterogeneous LLM Workflows
by: Huang, Zirui, et al.
Published: (2026)
by: Huang, Zirui, et al.
Published: (2026)
Barrier-Augmented Lagrangian for GPU-based Elastodynamic Contact
by: Guo, Dewen, et al.
Published: (2024)
by: Guo, Dewen, et al.
Published: (2024)
Accelerating Biclique Counting on GPU
by: Qiu, Linshan, et al.
Published: (2024)
by: Qiu, Linshan, et al.
Published: (2024)
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
by: Huang, En-Ming, et al.
Published: (2025)
by: Huang, En-Ming, et al.
Published: (2025)
HoSZp: An Efficient Homomorphic Error-bounded Lossy Compressor for Scientific Data
by: Agarwal, Tripti, et al.
Published: (2024)
by: Agarwal, Tripti, et al.
Published: (2024)
Efficient Graph Embedding at Scale: Optimizing CPU-GPU-SSD Integration
by: Li, Zhonggen, et al.
Published: (2025)
by: Li, Zhonggen, et al.
Published: (2025)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
by: Lin, Mao, et al.
Published: (2026)
by: Lin, Mao, et al.
Published: (2026)
Enabling mixed-precision in spectral element codes
by: Chen, Yanxiang, et al.
Published: (2025)
by: Chen, Yanxiang, et al.
Published: (2025)
Towards Fast Setup and High Throughput of GPU Serverless Computing
by: Zhao, Han, et al.
Published: (2024)
by: Zhao, Han, et al.
Published: (2024)
Hummingbird: SLO-Oriented GPU Preemption at Microsecond-scale
by: Hu, Tiancheng, et al.
Published: (2026)
by: Hu, Tiancheng, et al.
Published: (2026)
Heimdall++: Optimizing GPU Utilization and Pipeline Parallelism for Efficient Single-Pulse Detection
by: Xia, Bingzheng, et al.
Published: (2025)
by: Xia, Bingzheng, et al.
Published: (2025)
High-Performance N-Queens Solver on GPU: Iterative DFS with Zero Bank Conflicts
by: Yao, Guangchao, et al.
Published: (2025)
by: Yao, Guangchao, et al.
Published: (2025)
GPU-Accelerated Selected Basis Diagonalization with Thrust for SQD-based Algorithms
by: Doi, Jun, et al.
Published: (2026)
by: Doi, Jun, et al.
Published: (2026)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
by: Zhang, WenZheng, et al.
Published: (2024)
by: Zhang, WenZheng, et al.
Published: (2024)
GRNND: A GPU-Parallel Relative NN-Descent Algorithm for Efficient Approximate Nearest Neighbor Graph Construction
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
cuFastTuckerPlus: A Stochastic Parallel Sparse FastTucker Decomposition Using GPU Tensor Cores
by: Li, Zixuan, et al.
Published: (2024)
by: Li, Zixuan, et al.
Published: (2024)
SOLANET: Distributed Neighbor Graph Construction on GPU-Accelerated Systems
by: Iwabuchi, Keita, et al.
Published: (2026)
by: Iwabuchi, Keita, et al.
Published: (2026)
FlowWalker: A Memory-efficient and High-performance GPU-based Dynamic Graph Random Walk Framework
by: Mei, Junyi, et al.
Published: (2024)
by: Mei, Junyi, et al.
Published: (2024)
Cronus: Efficient LLM inference on Heterogeneous GPU Clusters via Partially Disaggregated Prefill
by: Liu, Yunzhao, et al.
Published: (2025)
by: Liu, Yunzhao, et al.
Published: (2025)
Multi-core & GPU-based Balanced Butterfly Counting in Signed Bipartite Graphs
by: Kiran, Mekala, et al.
Published: (2026)
by: Kiran, Mekala, et al.
Published: (2026)
An AD based library for Efficient Hessian and Hessian-Vector Product Computation on GPU
by: Ranjan, Desh, et al.
Published: (2024)
by: Ranjan, Desh, et al.
Published: (2024)
PRISM: Dynamic Primitive-Based Forecasting for Large-Scale GPU Cluster Workloads
by: Wu, Xin, et al.
Published: (2026)
by: Wu, Xin, et al.
Published: (2026)
APEX: Asynchronous Parallel CPU-GPU Execution for Online LLM Inference on Constrained GPUs
by: Fan, Jiakun, et al.
Published: (2025)
by: Fan, Jiakun, et al.
Published: (2025)
Neutron particle transport 3D method of characteristic Multi GPU platform Parallel Computing
by: Zhou, Faguo, et al.
Published: (2025)
by: Zhou, Faguo, et al.
Published: (2025)
HARP: Orchestrating Automated Parallel Training on Heterogeneous GPU Clusters
by: Liang, Antian, et al.
Published: (2025)
by: Liang, Antian, et al.
Published: (2025)
Similar Items
-
SIMT/GPU Data Race Verification using ISCC and Intermediary Code Representations: A Case Study
by: Osterhout, Andrew, et al.
Published: (2025) -
HiRace: Accurate and Fast Source-Level Race Checking of GPU Programs
by: Jacobson, John, et al.
Published: (2024) -
GPU-Accelerated Distributed QAOA on Large-scale HPC Ecosystems
by: Xu, Zhihao, et al.
Published: (2025) -
GPZ: GPU-Accelerated Lossy Compressor for Particle Data
by: Li, Ruoyu, et al.
Published: (2025) -
The Shamrock code: I- Smoothed Particle Hydrodynamics on GPUs
by: David--Cléris, Timothée, et al.
Published: (2025)