GPU acceleration of non-equilibrium Green's function calculation using OpenACC and CUDA FORTRAN
Fuente:
arXiv
Saved in:
| Main Authors: | Yin, Jia, Ibrahim, Khaled Z., Del Ben, Mauro, Deslippe, Jack, Chan, Yang-hao, Yang, Chao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GPU Acceleration of Learning With Errors KEMs Using OpenACC for Post-Quantum Cryptography
by: Liberati, Tiziana, et al.
Published: (2026)
by: Liberati, Tiziana, et al.
Published: (2026)
Accelerating the Particle-In-Cell code ECsim with OpenACC
by: Boella, Elisabetta, et al.
Published: (2026)
by: Boella, Elisabetta, et al.
Published: (2026)
Accelerating Particle-in-Cell Monte Carlo Simulations with MPI, OpenMP/OpenACC and Asynchronous Multi-GPU Programming
by: Williams, Jeremy J., et al.
Published: (2024)
by: Williams, Jeremy J., et al.
Published: (2024)
Alya towards Exascale: Optimal OpenACC Performance of the Navier-Stokes Finite Element Assembly on GPUs
by: Owen, Herbert, et al.
Published: (2024)
by: Owen, Herbert, et al.
Published: (2024)
GPU-Accelerated Algorithms for Process Mapping
by: Samoldekin, Petr, et al.
Published: (2025)
by: Samoldekin, Petr, et al.
Published: (2025)
Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor Cores
by: Zhao, Haisha, et al.
Published: (2025)
by: Zhao, Haisha, et al.
Published: (2025)
Bridging the Gap: Empowering Small Models in Reliable OpenACC-based Parallelization via GEPA-Optimized Prompting
by: Jhaveri, Samyak, et al.
Published: (2026)
by: Jhaveri, Samyak, et al.
Published: (2026)
Accelerating Intra-Node GPU-to-GPU Communication Through Multi-Path Transfers with CUDA Graphs
by: Sojoodi, Amirhossein, et al.
Published: (2026)
by: Sojoodi, Amirhossein, et al.
Published: (2026)
ACC Saturator: Automatic Kernel Optimization for Directive-Based GPU Code
by: Matsumura, Kazuaki, et al.
Published: (2023)
by: Matsumura, Kazuaki, et al.
Published: (2023)
Accelerating the Dutch Atmospheric Large-Eddy Simulation (DALES) model with OpenACC
by: Esclapez, Lucas, et al.
Published: (2025)
by: Esclapez, Lucas, et al.
Published: (2025)
OPTIMUM-DERAM: Highly Consistent, Scalable, and Secure Multi-Object Memory using RLNC
by: Nicolaou, Nicolas, et al.
Published: (2026)
by: Nicolaou, Nicolas, et al.
Published: (2026)
Context Adaptive Cooperation
by: Albouy, Timothé, et al.
Published: (2023)
by: Albouy, Timothé, et al.
Published: (2023)
Population Protocols Revisited: Parity and Beyond
by: Gąsieniec, Leszek, et al.
Published: (2025)
by: Gąsieniec, Leszek, et al.
Published: (2025)
Flexible Multi-Dimensional FFTs for Plane Wave Density Functional Theory Codes
by: Popovici, Doru Thom, et al.
Published: (2024)
by: Popovici, Doru Thom, et al.
Published: (2024)
Cost-Effective Methodology for Complex Tuning Searches in HPC: Navigating Interdependencies and Dimensionality
by: Dieguez, Adrian Perez, et al.
Published: (2024)
by: Dieguez, Adrian Perez, et al.
Published: (2024)
Parallel Paradigms in Modern HPC: A Comparative Analysis of MPI, OpenMP, and CUDA
by: ALHafez, Nizar, et al.
Published: (2025)
by: ALHafez, Nizar, et al.
Published: (2025)
Combining GPU and CPU for accelerating evolutionary computing workloads
by: Eynaliyev, Rustam, et al.
Published: (2025)
by: Eynaliyev, Rustam, et al.
Published: (2025)
A GPU-accelerated Molecular Docking Workflow with Kubernetes and Apache Airflow
by: Medeiros, Daniel, et al.
Published: (2024)
by: Medeiros, Daniel, et al.
Published: (2024)
Parallel Gaussian process with kernel approximation in CUDA
by: Carminati, Davide
Published: (2024)
by: Carminati, Davide
Published: (2024)
A Study of Performance Programming of CPU, GPU accelerated Computers and SIMD Architecture
by: Yi, Xinyao
Published: (2024)
by: Yi, Xinyao
Published: (2024)
A Morton-Type Space-Filling Curve for Pyramid Subdivision and Hybrid Adaptive Mesh Refinement
by: Knapp, David, et al.
Published: (2026)
by: Knapp, David, et al.
Published: (2026)
WgPy: GPU-accelerated NumPy-like array library for web browsers
by: Hidaka, Masatoshi, et al.
Published: (2025)
by: Hidaka, Masatoshi, et al.
Published: (2025)
High-Performance Parallelization of Dijkstra's Algorithm Using MPI and CUDA
by: Song, Boyang
Published: (2025)
by: Song, Boyang
Published: (2025)
Distributed OpenMP Offloading of OpenMC on Intel GPU MAX Accelerators
by: Fridman, Yehonatan, et al.
Published: (2024)
by: Fridman, Yehonatan, et al.
Published: (2024)
Parallel DNA Sequence Alignment on High-Performance Systems with CUDA and MPI
by: Zwaka, Linus
Published: (2024)
by: Zwaka, Linus
Published: (2024)
Boosting Performance of Iterative Applications on GPUs: Kernel Batching with CUDA Graphs
by: Ekelund, Jonah, et al.
Published: (2025)
by: Ekelund, Jonah, et al.
Published: (2025)
A GPU accelerated mixed-precision Smoothed Particle Hydrodynamics framework with cell-based relative coordinates
by: Mao, Zirui, et al.
Published: (2023)
by: Mao, Zirui, et al.
Published: (2023)
Debunking the CUDA Myth Towards GPU-based AI Systems
by: Lee, Yunjae, et al.
Published: (2024)
by: Lee, Yunjae, et al.
Published: (2024)
Intel(R) SHMEM: GPU-initiated OpenSHMEM using SYCL
by: Brooks, Alex, et al.
Published: (2024)
by: Brooks, Alex, et al.
Published: (2024)
Lessons Learned Migrating CUDA to SYCL: A HEP Case Study with ROOT RDataFrame
by: Chen, Jolly, et al.
Published: (2024)
by: Chen, Jolly, et al.
Published: (2024)
GPU-Accelerated Batch-Dynamic Subgraph Matching
by: Qiu, Linshan, et al.
Published: (2024)
by: Qiu, Linshan, et al.
Published: (2024)
AGILE: Lightweight and Efficient Asynchronous GPU-SSD Integration
by: Yang, Zhuoping, et al.
Published: (2025)
by: Yang, Zhuoping, et al.
Published: (2025)
FlashSpread: IO-Aware GPU Simulation of Non-Markovian Epidemic Dynamics via Kernel Fusion
by: Shakeri, Heman, et al.
Published: (2026)
by: Shakeri, Heman, et al.
Published: (2026)
Dynamic Approximate Maximum Matching in the Distributed Vertex Partition Model
by: Robinson, Peter, et al.
Published: (2025)
by: Robinson, Peter, et al.
Published: (2025)
Dynamic Memory Management on GPUs with SYCL
by: Standish, Russell K.
Published: (2025)
by: Standish, Russell K.
Published: (2025)
Data Scheduling Algorithm for Scalable and Efficient IoT Sensing in Cloud Computing
by: Mohammad, Noor Islam S.
Published: (2025)
by: Mohammad, Noor Islam S.
Published: (2025)
Distributed Tomographic Reconstruction with Quantization
by: Miao, Runxuan, et al.
Published: (2024)
by: Miao, Runxuan, et al.
Published: (2024)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
by: Zhang, WenZheng, et al.
Published: (2024)
by: Zhang, WenZheng, et al.
Published: (2024)
Secure Federated XGBoost with CUDA-accelerated Homomorphic Encryption via NVIDIA FLARE
by: Xu, Ziyue, et al.
Published: (2025)
by: Xu, Ziyue, et al.
Published: (2025)
Beyond Microservices: Testing Web-Scale RCA Methods on GPU-Driven LLM Workloads
by: Scheinert, Dominik, et al.
Published: (2026)
by: Scheinert, Dominik, et al.
Published: (2026)
Similar Items
-
GPU Acceleration of Learning With Errors KEMs Using OpenACC for Post-Quantum Cryptography
by: Liberati, Tiziana, et al.
Published: (2026) -
Accelerating the Particle-In-Cell code ECsim with OpenACC
by: Boella, Elisabetta, et al.
Published: (2026) -
Accelerating Particle-in-Cell Monte Carlo Simulations with MPI, OpenMP/OpenACC and Asynchronous Multi-GPU Programming
by: Williams, Jeremy J., et al.
Published: (2024) -
Alya towards Exascale: Optimal OpenACC Performance of the Navier-Stokes Finite Element Assembly on GPUs
by: Owen, Herbert, et al.
Published: (2024) -
GPU-Accelerated Algorithms for Process Mapping
by: Samoldekin, Petr, et al.
Published: (2025)