A Precision Emulation Approach to the GPU Acceleration of Ab Initio Electronic Structure Calculations
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Hang, Li, Junjie, Wang, Yinzhi, Nepal, Niraj K., Wang, Yang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Pilot Study on Tunable Precision Emulation via Automatic BLAS Offloading
por: Liu, Hang, et al.
Publicado: (2025)
por: Liu, Hang, et al.
Publicado: (2025)
Fast and Scalable Mixed Precision Euclidean Distance Calculations Using GPU Tensor Cores
por: Curless, Brian, et al.
Publicado: (2025)
por: Curless, Brian, et al.
Publicado: (2025)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
por: Zhang, Li, et al.
Publicado: (2025)
por: Zhang, Li, et al.
Publicado: (2025)
Disaggregated Design for GPU-Based Volumetric Data Structures
por: Meneghin, Massimiliano, et al.
Publicado: (2025)
por: Meneghin, Massimiliano, et al.
Publicado: (2025)
Unleashing the Power of Preemptive Priority-based Scheduling for Real-Time GPU Tasks
por: Wang, Yidi, et al.
Publicado: (2024)
por: Wang, Yidi, et al.
Publicado: (2024)
CUTHERMO: Understanding GPU Memory Inefficiencies with Heat Map Profiling
por: Zhao, Yanbo, et al.
Publicado: (2025)
por: Zhao, Yanbo, et al.
Publicado: (2025)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
por: Lin, Mao, et al.
Publicado: (2026)
por: Lin, Mao, et al.
Publicado: (2026)
Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
por: Liu, Shifang, et al.
Publicado: (2025)
por: Liu, Shifang, et al.
Publicado: (2025)
Toward Scalable Docker-Based Emulations of Blockchain Networks for Research and Development
por: Pennino, Diego, et al.
Publicado: (2024)
por: Pennino, Diego, et al.
Publicado: (2024)
The Energy Cost of Execution-Idle in GPU Clusters
por: Lei, Yiran, et al.
Publicado: (2026)
por: Lei, Yiran, et al.
Publicado: (2026)
Scalable GPU Performance Variability Analysis framework
por: Lahiry, Ankur, et al.
Publicado: (2025)
por: Lahiry, Ankur, et al.
Publicado: (2025)
On the Partitioning of GPU Power among Multi-Instances
por: Vamja, Tirth, et al.
Publicado: (2025)
por: Vamja, Tirth, et al.
Publicado: (2025)
Taking GPU Programming Models to Task for Performance Portability
por: Davis, Joshua H., et al.
Publicado: (2024)
por: Davis, Joshua H., et al.
Publicado: (2024)
Profiling and optimization of multi-card GPU machine learning jobs
por: Lawenda, Marcin, et al.
Publicado: (2025)
por: Lawenda, Marcin, et al.
Publicado: (2025)
KEET: Explaining Performance of GPU Kernels Using LLM Agents
por: Davis, Joshua H., et al.
Publicado: (2026)
por: Davis, Joshua H., et al.
Publicado: (2026)
High-Performance Portable GPU Primitives for Arbitrary Types and Operators in Julia
por: Pilliat, Emmanuel
Publicado: (2026)
por: Pilliat, Emmanuel
Publicado: (2026)
Efficient allocation of image recognition and LLM tasks on multi-GPU system
por: Lawenda, Marcin, et al.
Publicado: (2025)
por: Lawenda, Marcin, et al.
Publicado: (2025)
Data-Driven Analysis to Understand GPU Hardware Resource Usage of Optimizations
por: Islam, Tanzima Z., et al.
Publicado: (2024)
por: Islam, Tanzima Z., et al.
Publicado: (2024)
Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters
por: Jain, Rutwik, et al.
Publicado: (2026)
por: Jain, Rutwik, et al.
Publicado: (2026)
Automatic BLAS Offloading on Unified Memory Architecture: A Study on NVIDIA Grace-Hopper
por: Li, Junjie, et al.
Publicado: (2024)
por: Li, Junjie, et al.
Publicado: (2024)
Accelerating Particle-in-Cell Monte Carlo Simulations with MPI, OpenMP/OpenACC and Asynchronous Multi-GPU Programming
por: Williams, Jeremy J., et al.
Publicado: (2024)
por: Williams, Jeremy J., et al.
Publicado: (2024)
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
por: Wahlgren, Jacob, et al.
Publicado: (2025)
por: Wahlgren, Jacob, et al.
Publicado: (2025)
LEO: Tracing GPU Stall Root Causes via Cross-Vendor Backward Slicing
por: Xia, Yuning, et al.
Publicado: (2026)
por: Xia, Yuning, et al.
Publicado: (2026)
Accelerating Mobile Inference through Fine-Grained CPU-GPU Co-Execution
por: Li, Zhuojin, et al.
Publicado: (2025)
por: Li, Zhuojin, et al.
Publicado: (2025)
Towards Portability at Scale: A Cross-Architecture Performance Evaluation of a GPU-enabled Shallow Water Solver
por: Villalobos, Johansell, et al.
Publicado: (2025)
por: Villalobos, Johansell, et al.
Publicado: (2025)
PASTA: A Modular Program Analysis Tool Framework for Accelerators
por: Lin, Mao, et al.
Publicado: (2026)
por: Lin, Mao, et al.
Publicado: (2026)
Accelerating Gaussian beam tracing method with dynamic parallelism on graphics processing units
por: Sheng, Zhang, et al.
Publicado: (2025)
por: Sheng, Zhang, et al.
Publicado: (2025)
Hardware-Agnostic and Insightful Efficiency Metrics for Accelerated Systems: Definition and Implementation within TALP
por: Rahimi, Ghazal, et al.
Publicado: (2026)
por: Rahimi, Ghazal, et al.
Publicado: (2026)
AcceleratedKernels.jl: Cross-Architecture Parallel Algorithms from a Unified, Transpiled Codebase
por: Nicusan, Andrei-Leonard, et al.
Publicado: (2025)
por: Nicusan, Andrei-Leonard, et al.
Publicado: (2025)
Orthrus: Accelerating Multi-BFT Consensus through Concurrent Partial Ordering of Transactions (Extended Version)
por: Lyu, Hanzheng, et al.
Publicado: (2024)
por: Lyu, Hanzheng, et al.
Publicado: (2024)
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
por: Wang, Yuxin, et al.
Publicado: (2024)
por: Wang, Yuxin, et al.
Publicado: (2024)
Matryoshka: Optimization of Dynamic Diverse Quantum Chemistry Systems via Elastic Parallelism Transformation
por: Wang, Tuowei, et al.
Publicado: (2024)
por: Wang, Tuowei, et al.
Publicado: (2024)
Staging Blocked Evaluation over Structured Sparse Matrices
por: Das, Pratyush, et al.
Publicado: (2024)
por: Das, Pratyush, et al.
Publicado: (2024)
mLR: Scalable Laminography Reconstruction based on Memoization
por: Ma, Bin, et al.
Publicado: (2025)
por: Ma, Bin, et al.
Publicado: (2025)
Redesigning GROMACS Halo Exchange: Improving Strong Scaling with GPU-initiated NVSHMEM
por: Doijade, Mahesh, et al.
Publicado: (2025)
por: Doijade, Mahesh, et al.
Publicado: (2025)
Arkade: k-Nearest Neighbor Search With Non-Euclidean Distances using GPU Ray Tracing
por: Mandarapu, Durga, et al.
Publicado: (2023)
por: Mandarapu, Durga, et al.
Publicado: (2023)
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
por: Zhang, Yaozheng, et al.
Publicado: (2025)
por: Zhang, Yaozheng, et al.
Publicado: (2025)
The Landscape of GPU-Centric Communication
por: Unat, Didem, et al.
Publicado: (2024)
por: Unat, Didem, et al.
Publicado: (2024)
GigaAPI for GPU Parallelization
por: Suvarna, M., et al.
Publicado: (2025)
por: Suvarna, M., et al.
Publicado: (2025)
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
por: Zhao, Xuanlei, et al.
Publicado: (2024)
por: Zhao, Xuanlei, et al.
Publicado: (2024)
Ejemplares similares
-
A Pilot Study on Tunable Precision Emulation via Automatic BLAS Offloading
por: Liu, Hang, et al.
Publicado: (2025) -
Fast and Scalable Mixed Precision Euclidean Distance Calculations Using GPU Tensor Cores
por: Curless, Brian, et al.
Publicado: (2025) -
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
por: Zhang, Li, et al.
Publicado: (2025) -
Disaggregated Design for GPU-Based Volumetric Data Structures
por: Meneghin, Massimiliano, et al.
Publicado: (2025) -
Unleashing the Power of Preemptive Priority-based Scheduling for Real-Time GPU Tasks
por: Wang, Yidi, et al.
Publicado: (2024)