A Pilot Study on Tunable Precision Emulation via Automatic BLAS Offloading
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Hang, Li, Junjie, Wang, Yinzhi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Automatic BLAS Offloading on Unified Memory Architecture: A Study on NVIDIA Grace-Hopper
von: Li, Junjie, et al.
Veröffentlicht: (2024)
von: Li, Junjie, et al.
Veröffentlicht: (2024)
A Precision Emulation Approach to the GPU Acceleration of Ab Initio Electronic Structure Calculations
von: Liu, Hang, et al.
Veröffentlicht: (2026)
von: Liu, Hang, et al.
Veröffentlicht: (2026)
Performant Automatic BLAS Offloading on Unified Memory Architecture with OpenMP First-Touch Style Data Movement
von: Li, Junjie
Veröffentlicht: (2024)
von: Li, Junjie
Veröffentlicht: (2024)
Performance optimization of BLAS algorithms with band matrices for RISC-V processors
von: Pirova, Anna, et al.
Veröffentlicht: (2025)
von: Pirova, Anna, et al.
Veröffentlicht: (2025)
Toward Scalable Docker-Based Emulations of Blockchain Networks for Research and Development
von: Pennino, Diego, et al.
Veröffentlicht: (2024)
von: Pennino, Diego, et al.
Veröffentlicht: (2024)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
von: Zhang, Li, et al.
Veröffentlicht: (2025)
von: Zhang, Li, et al.
Veröffentlicht: (2025)
ADELIA: Automatic Differentiation for Efficient Laplace Inference Approximations
von: Boudaoud, Afif, et al.
Veröffentlicht: (2026)
von: Boudaoud, Afif, et al.
Veröffentlicht: (2026)
Fast and Scalable Mixed Precision Euclidean Distance Calculations Using GPU Tensor Cores
von: Curless, Brian, et al.
Veröffentlicht: (2025)
von: Curless, Brian, et al.
Veröffentlicht: (2025)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
von: Lin, Mao, et al.
Veröffentlicht: (2026)
von: Lin, Mao, et al.
Veröffentlicht: (2026)
Cloud Performance Decomposition for Long-Term Performance Engineering: A Case Study
von: Debnath, Shimul, et al.
Veröffentlicht: (2026)
von: Debnath, Shimul, et al.
Veröffentlicht: (2026)
An Experimental Study of Different Aggregation Schemes in Semi-Asynchronous Federated Learning
von: Li, Yunbo, et al.
Veröffentlicht: (2024)
von: Li, Yunbo, et al.
Veröffentlicht: (2024)
Matryoshka: Optimization of Dynamic Diverse Quantum Chemistry Systems via Elastic Parallelism Transformation
von: Wang, Tuowei, et al.
Veröffentlicht: (2024)
von: Wang, Tuowei, et al.
Veröffentlicht: (2024)
WebAssembly and Unikernels: A Comparative Study for Serverless at the Edge
von: Besozzi, Valerio, et al.
Veröffentlicht: (2025)
von: Besozzi, Valerio, et al.
Veröffentlicht: (2025)
Automated Calibration of Parallel and Distributed Computing Simulators: A Case Study
von: McDonald, Jesse, et al.
Veröffentlicht: (2024)
von: McDonald, Jesse, et al.
Veröffentlicht: (2024)
CloverLeaf on Intel Multi-Core CPUs: A Case Study in Write-Allocate Evasion
von: Laukemann, Jan, et al.
Veröffentlicht: (2023)
von: Laukemann, Jan, et al.
Veröffentlicht: (2023)
DIAL: Decentralized I/O AutoTuning via Learned Client-side Local Metrics for Parallel File System
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
CUTHERMO: Understanding GPU Memory Inefficiencies with Heat Map Profiling
von: Zhao, Yanbo, et al.
Veröffentlicht: (2025)
von: Zhao, Yanbo, et al.
Veröffentlicht: (2025)
Unleashing the Power of Preemptive Priority-based Scheduling for Real-Time GPU Tasks
von: Wang, Yidi, et al.
Veröffentlicht: (2024)
von: Wang, Yidi, et al.
Veröffentlicht: (2024)
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
von: Wang, Yuxin, et al.
Veröffentlicht: (2024)
von: Wang, Yuxin, et al.
Veröffentlicht: (2024)
A dynamic parallel method for performance optimization on hybrid CPUs
von: Yu, Luo, et al.
Veröffentlicht: (2024)
von: Yu, Luo, et al.
Veröffentlicht: (2024)
Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
von: Liu, Shifang, et al.
Veröffentlicht: (2025)
von: Liu, Shifang, et al.
Veröffentlicht: (2025)
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
von: Jeong, Bodon, et al.
Veröffentlicht: (2026)
von: Jeong, Bodon, et al.
Veröffentlicht: (2026)
mLR: Scalable Laminography Reconstruction based on Memoization
von: Ma, Bin, et al.
Veröffentlicht: (2025)
von: Ma, Bin, et al.
Veröffentlicht: (2025)
LEO: Tracing GPU Stall Root Causes via Cross-Vendor Backward Slicing
von: Xia, Yuning, et al.
Veröffentlicht: (2026)
von: Xia, Yuning, et al.
Veröffentlicht: (2026)
"Two-Stagification": Job Dispatching in Large-Scale Clusters via a Two-Stage Architecture
von: Yildiz, Mert, et al.
Veröffentlicht: (2025)
von: Yildiz, Mert, et al.
Veröffentlicht: (2025)
THEAS: Efficient Power Management in Multi-Core CPUs via Cache-Aware Resource Scheduling
von: Muhammad, Said, et al.
Veröffentlicht: (2025)
von: Muhammad, Said, et al.
Veröffentlicht: (2025)
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
von: Zhang, Yaozheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yaozheng, et al.
Veröffentlicht: (2025)
Parallel I/O Characterization and Optimization on Large-Scale HPC Systems: A 360-Degree Survey
von: Ather, Hammad, et al.
Veröffentlicht: (2024)
von: Ather, Hammad, et al.
Veröffentlicht: (2024)
Confidential Computing on NVIDIA Hopper GPUs: A Performance Benchmark Study
von: Zhu, Jianwei, et al.
Veröffentlicht: (2024)
von: Zhu, Jianwei, et al.
Veröffentlicht: (2024)
Optimal Parallel Scheduling under Concave Speedup Functions
von: Li, Chengzhang, et al.
Veröffentlicht: (2025)
von: Li, Chengzhang, et al.
Veröffentlicht: (2025)
Resource Management Schemes for Cloud-Native Platforms with Computing Containers of Docker and Kubernetes
von: Mao, Ying, et al.
Veröffentlicht: (2020)
von: Mao, Ying, et al.
Veröffentlicht: (2020)
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
How to Rent GPUs on a Budget
von: Li, Zhouzi, et al.
Veröffentlicht: (2024)
von: Li, Zhouzi, et al.
Veröffentlicht: (2024)
Asymptotically Optimal Scheduling of Multiple Parallelizable Job Classes
von: Berg, Benjamin, et al.
Veröffentlicht: (2024)
von: Berg, Benjamin, et al.
Veröffentlicht: (2024)
Recorder: Comprehensive Parallel I/O Tracing and Analysis
von: Wang, Chen, et al.
Veröffentlicht: (2025)
von: Wang, Chen, et al.
Veröffentlicht: (2025)
Shifting the Sweet Spot: High-Performance Matrix-Free Method for High-Order Elasticity
von: Chang, Dali, et al.
Veröffentlicht: (2026)
von: Chang, Dali, et al.
Veröffentlicht: (2026)
An Online Probabilistic Distributed Tracing System
von: Toslali, M., et al.
Veröffentlicht: (2024)
von: Toslali, M., et al.
Veröffentlicht: (2024)
A Performance Analysis of BFT Consensus for Blockchains
von: Chan, J. D., et al.
Veröffentlicht: (2024)
von: Chan, J. D., et al.
Veröffentlicht: (2024)
Sampling in Cloud Benchmarking: A Critical Review and Methodological Guidelines
von: Akbari, Saman, et al.
Veröffentlicht: (2025)
von: Akbari, Saman, et al.
Veröffentlicht: (2025)
PASTA: A Modular Program Analysis Tool Framework for Accelerators
von: Lin, Mao, et al.
Veröffentlicht: (2026)
von: Lin, Mao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Automatic BLAS Offloading on Unified Memory Architecture: A Study on NVIDIA Grace-Hopper
von: Li, Junjie, et al.
Veröffentlicht: (2024) -
A Precision Emulation Approach to the GPU Acceleration of Ab Initio Electronic Structure Calculations
von: Liu, Hang, et al.
Veröffentlicht: (2026) -
Performant Automatic BLAS Offloading on Unified Memory Architecture with OpenMP First-Touch Style Data Movement
von: Li, Junjie
Veröffentlicht: (2024) -
Performance optimization of BLAS algorithms with band matrices for RISC-V processors
von: Pirova, Anna, et al.
Veröffentlicht: (2025) -
Toward Scalable Docker-Based Emulations of Blockchain Networks for Research and Development
von: Pennino, Diego, et al.
Veröffentlicht: (2024)