Gespeichert in:
| Hauptverfasser: | Liu, Rongrong, Guo, Zhuoqiang, Sha, Qiuchen, Zhao, Tong, Li, Haibo, Hu, Wei, Liu, Lijun, Tan, Guangming, Jia, Weile |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2501.03061 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Deep Learning-Enabled Supercritical Flame Simulation at Detailed Chemistry and Real-Fluid Accuracy Towards Trillion-Cell Scale
von: Guo, Zhuoqiang, et al.
Veröffentlicht: (2025)
von: Guo, Zhuoqiang, et al.
Veröffentlicht: (2025)
JanusPipe: Efficient Pipeline Parallel Training for Machine Learning Interatomic Potentials
von: Wang, Hongyu, et al.
Veröffentlicht: (2026)
von: Wang, Hongyu, et al.
Veröffentlicht: (2026)
Breaking the Training Barrier of Billion-Parameter Universal Machine Learning Interatomic Potentials
von: Zhou, Yuanchang, et al.
Veröffentlicht: (2026)
von: Zhou, Yuanchang, et al.
Veröffentlicht: (2026)
Scaling Molecular Dynamics with ab initio Accuracy to 149 Nanoseconds per Day
von: Li, Jianxiong, et al.
Veröffentlicht: (2024)
von: Li, Jianxiong, et al.
Veröffentlicht: (2024)
FastCHGNet: Training one Universal Interatomic Potential to 1.5 Hours with 32 GPUs
von: Zhou, Yuanchang, et al.
Veröffentlicht: (2024)
von: Zhou, Yuanchang, et al.
Veröffentlicht: (2024)
Multi-level Memory-Centric Profiling on ARM Processors with ARM SPE
von: Miksits, Samuel, et al.
Veröffentlicht: (2024)
von: Miksits, Samuel, et al.
Veröffentlicht: (2024)
GPU-Accelerated Modified Bessel Function of the Second Kind for Gaussian Processes
von: Geng, Zipei, et al.
Veröffentlicht: (2025)
von: Geng, Zipei, et al.
Veröffentlicht: (2025)
Pipelined Dense Symmetric Eigenvalue Decomposition on Multi-GPU Architectures
von: Wang, Hansheng, et al.
Veröffentlicht: (2025)
von: Wang, Hansheng, et al.
Veröffentlicht: (2025)
GCAPS: GPU Context-Aware Preemptive Priority-based Scheduling for Real-Time Tasks
von: Wang, Yidi, et al.
Veröffentlicht: (2024)
von: Wang, Yidi, et al.
Veröffentlicht: (2024)
Large-scale Neural Network Quantum States for ab initio Quantum Chemistry Simulations on Fugaku
von: Xu, Hongtao, et al.
Veröffentlicht: (2025)
von: Xu, Hongtao, et al.
Veröffentlicht: (2025)
Exploring the Viability of Unikernels for ARM-powered Edge Computing
von: Kaiser, Shahidullah, et al.
Veröffentlicht: (2024)
von: Kaiser, Shahidullah, et al.
Veröffentlicht: (2024)
Demystifying ARM SME to Optimize General Matrix Multiplications
von: Deng, Chencheng, et al.
Veröffentlicht: (2025)
von: Deng, Chencheng, et al.
Veröffentlicht: (2025)
Dependency-aware Resource Allocation for Serverless Functions at the Edge
von: Baresi, Luciano, et al.
Veröffentlicht: (2023)
von: Baresi, Luciano, et al.
Veröffentlicht: (2023)
A Framework for Fine-Grained Synchronization of Dependent GPU Kernels
von: Jangda, Abhinav, et al.
Veröffentlicht: (2023)
von: Jangda, Abhinav, et al.
Veröffentlicht: (2023)
Orchestrating the Execution of Serverless Functions in Hybrid Clouds
von: Peri, Aristotelis, et al.
Veröffentlicht: (2024)
von: Peri, Aristotelis, et al.
Veröffentlicht: (2024)
Computational Performance and Energy Efficiency of ARM based HPC servers
von: Schirmer, Oskar
Veröffentlicht: (2024)
von: Schirmer, Oskar
Veröffentlicht: (2024)
A Hybrid Vectorized Merge Sort on ARM NEON
von: Zhou, Jincheng, et al.
Veröffentlicht: (2024)
von: Zhou, Jincheng, et al.
Veröffentlicht: (2024)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
von: Gu, Jianfeng, et al.
Veröffentlicht: (2025)
von: Gu, Jianfeng, et al.
Veröffentlicht: (2025)
Unleashing the Power of Preemptive Priority-based Scheduling for Real-Time GPU Tasks
von: Wang, Yidi, et al.
Veröffentlicht: (2024)
von: Wang, Yidi, et al.
Veröffentlicht: (2024)
MQFQ-Sticky: Fair Queueing For Serverless GPU Functions
von: Fuerst, Alexander, et al.
Veröffentlicht: (2025)
von: Fuerst, Alexander, et al.
Veröffentlicht: (2025)
Heat: Satellite's meat is GPU's poison
von: Yuan, Zhehu, et al.
Veröffentlicht: (2024)
von: Yuan, Zhehu, et al.
Veröffentlicht: (2024)
Combining GPU and CPU for accelerating evolutionary computing workloads
von: Eynaliyev, Rustam, et al.
Veröffentlicht: (2025)
von: Eynaliyev, Rustam, et al.
Veröffentlicht: (2025)
ARM SVE Unleashed: Performance and Insights Across HPC Applications on Nvidia Grace
von: Shi, Ruimin, et al.
Veröffentlicht: (2025)
von: Shi, Ruimin, et al.
Veröffentlicht: (2025)
High-performance Vector-length Agnostic Quantum Circuit Simulations on ARM Processors
von: Shi, Ruimin, et al.
Veröffentlicht: (2026)
von: Shi, Ruimin, et al.
Veröffentlicht: (2026)
Breaking the Memory Wall: A Study of I/O Patterns and GPU Memory Utilization for Hybrid CPU-GPU Offloaded Optimizers
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
Fast and Scalable Mixed Precision Euclidean Distance Calculations Using GPU Tensor Cores
von: Curless, Brian, et al.
Veröffentlicht: (2025)
von: Curless, Brian, et al.
Veröffentlicht: (2025)
Characterization-Guided GPU Fault Resilience in NVIDIA MPS
von: Liu, Rixin, et al.
Veröffentlicht: (2026)
von: Liu, Rixin, et al.
Veröffentlicht: (2026)
Performance Isolation and Semantic Determinism in Efficient GPU Spatial Sharing
von: Yang, Zhenyuan, et al.
Veröffentlicht: (2026)
von: Yang, Zhenyuan, et al.
Veröffentlicht: (2026)
Parallel GPU-Enabled Algorithms for SpGEMM on Arbitrary Semirings with Hybrid Communication
von: McFarland, Thomas, et al.
Veröffentlicht: (2025)
von: McFarland, Thomas, et al.
Veröffentlicht: (2025)
Fantasy: Efficient Large-scale Vector Search on GPU Clusters with GPUDirect Async
von: Liu, Yi, et al.
Veröffentlicht: (2025)
von: Liu, Yi, et al.
Veröffentlicht: (2025)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
von: Mo, Zizhao, et al.
Veröffentlicht: (2026)
von: Mo, Zizhao, et al.
Veröffentlicht: (2026)
ICPS: Real-Time Resource Configuration for Cloud Serverless Functions Considering Affinity
von: Chen, Long, et al.
Veröffentlicht: (2025)
von: Chen, Long, et al.
Veröffentlicht: (2025)
A Precision Emulation Approach to the GPU Acceleration of Ab Initio Electronic Structure Calculations
von: Liu, Hang, et al.
Veröffentlicht: (2026)
von: Liu, Hang, et al.
Veröffentlicht: (2026)
SWIFT: Expedited Failure Recovery for Large-scale DNN Training
von: Zhong, Yuchen, et al.
Veröffentlicht: (2023)
von: Zhong, Yuchen, et al.
Veröffentlicht: (2023)
GPU-Accelerated Batch-Dynamic Subgraph Matching
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
SDSL-Solver: Scalable Distributed Sparse Linear Solvers for Large-Scale Interior Point Methods
von: Yang, Shaofeng, et al.
Veröffentlicht: (2026)
von: Yang, Shaofeng, et al.
Veröffentlicht: (2026)
AsyncSparse: Accelerating Sparse Matrix-Matrix Multiplication on Asynchronous GPU Architectures
von: Liu, Jie, et al.
Veröffentlicht: (2026)
von: Liu, Jie, et al.
Veröffentlicht: (2026)
A Preliminary Study on Accelerating Simulation Optimization with GPU Implementation
von: He, Jinghai, et al.
Veröffentlicht: (2024)
von: He, Jinghai, et al.
Veröffentlicht: (2024)
HC-SpMM: Accelerating Sparse Matrix-Matrix Multiplication for Graphs with Hybrid GPU Cores
von: Li, Zhonggen, et al.
Veröffentlicht: (2024)
von: Li, Zhonggen, et al.
Veröffentlicht: (2024)
ElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal Parallelism
von: Liu, Zedong, et al.
Veröffentlicht: (2025)
von: Liu, Zedong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Deep Learning-Enabled Supercritical Flame Simulation at Detailed Chemistry and Real-Fluid Accuracy Towards Trillion-Cell Scale
von: Guo, Zhuoqiang, et al.
Veröffentlicht: (2025) -
JanusPipe: Efficient Pipeline Parallel Training for Machine Learning Interatomic Potentials
von: Wang, Hongyu, et al.
Veröffentlicht: (2026) -
Breaking the Training Barrier of Billion-Parameter Universal Machine Learning Interatomic Potentials
von: Zhou, Yuanchang, et al.
Veröffentlicht: (2026) -
Scaling Molecular Dynamics with ab initio Accuracy to 149 Nanoseconds per Day
von: Li, Jianxiong, et al.
Veröffentlicht: (2024) -
FastCHGNet: Training one Universal Interatomic Potential to 1.5 Hours with 32 GPUs
von: Zhou, Yuanchang, et al.
Veröffentlicht: (2024)