Profiling Apple Silicon Performance for ML Training
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Dahua, Xu, Zhiming, Wang, Rongxiang, Lin, Felix Xiaozhu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Profiling Large Language Model Inference on Apple Silicon: A Quantization Perspective
by: Benazir, Afsara, et al.
Published: (2025)
by: Benazir, Afsara, et al.
Published: (2025)
Accelerating Sparse Ternary GEMM for Quantized ML on Apple Silicon
by: Lipshitz, Baraq, et al.
Published: (2025)
by: Lipshitz, Baraq, et al.
Published: (2025)
RWKV-edge: Deeply Compressed RWKV for Resource-Constrained Devices
by: Choe, Wonkyo, et al.
Published: (2024)
by: Choe, Wonkyo, et al.
Published: (2024)
WhisperFlow: speech foundation models in real time
by: Wang, Rongxiang, et al.
Published: (2024)
by: Wang, Rongxiang, et al.
Published: (2024)
Evaluation of Domain-Specific Architectures for General-Purpose Applications in Apple Silicon
by: López, Álvaro Corrochano, et al.
Published: (2025)
by: López, Álvaro Corrochano, et al.
Published: (2025)
From 8 Seconds to 370ms: Kernel-Fused SAR Imaging on Apple Silicon via Single-Dispatch FFT Pipelines
by: Bergach, Mohamed Amine
Published: (2026)
by: Bergach, Mohamed Amine
Published: (2026)
When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon
by: Bergach, Mohamed Amine
Published: (2026)
by: Bergach, Mohamed Amine
Published: (2026)
Range, Not Precision: Block-Floating-Point Half-Precision FFT and SAR Imaging on Apple Silicon
by: Bergach, Mohamed Amine
Published: (2026)
by: Bergach, Mohamed Amine
Published: (2026)
Efficient Mixture-of-Experts LLM Inference with Apple Silicon NPUs
by: Benazir, Afsara, et al.
Published: (2026)
by: Benazir, Afsara, et al.
Published: (2026)
Performance Characterization and Optimizations of Traditional ML Applications
by: Kumar, Harsh, et al.
Published: (2024)
by: Kumar, Harsh, et al.
Published: (2024)
Turbocharge Speech Understanding with Pilot Inference
by: Wang, Rongxiang, et al.
Published: (2023)
by: Wang, Rongxiang, et al.
Published: (2023)
SysOM-AI: Continuous Cross-Layer Performance Diagnosis for Production AI Training
by: Zheng, Yusheng, et al.
Published: (2026)
by: Zheng, Yusheng, et al.
Published: (2026)
OSCAR-P and aMLLibrary: Profiling and Predicting the Performance of FaaS-based Applications in Computing Continua
by: Sala, Roberto, et al.
Published: (2024)
by: Sala, Roberto, et al.
Published: (2024)
Dissecting RISC-V Performance: Practical PMU Profiling and Hardware-Agnostic Roofline Analysis on Emerging Platforms
by: Batashev, Alexander
Published: (2025)
by: Batashev, Alexander
Published: (2025)
From Profiling to Optimization: Unveiling the Profile Guided Optimization
by: Liu, Bingxin, et al.
Published: (2025)
by: Liu, Bingxin, et al.
Published: (2025)
TrainMover: An Interruption-Resilient Runtime for ML Training
by: Lao, ChonLam, et al.
Published: (2024)
by: Lao, ChonLam, et al.
Published: (2024)
Scalable Packed Layouts for Vector-Length-Agnostic ML Code Generation
by: Beysel, Ege, et al.
Published: (2026)
by: Beysel, Ege, et al.
Published: (2026)
A Data-driven ML Approach for Maximizing Performance in LLM-Adapter Serving
by: Agullo, Ferran, et al.
Published: (2025)
by: Agullo, Ferran, et al.
Published: (2025)
gpu tracker: Python Package for Tracking and Profiling GPU and Other Hardware Utilization in Both Desktop and High-Performance Computing Environments
by: Huckvale, Erik D., et al.
Published: (2024)
by: Huckvale, Erik D., et al.
Published: (2024)
Towards Universal Performance Modeling for Machine Learning Training on Multi-GPU Platforms
by: Lin, Zhongyi, et al.
Published: (2024)
by: Lin, Zhongyi, et al.
Published: (2024)
Beating vDSP: A 138 GFLOPS Radix-8 Stockham FFT on Apple Silicon via Two-Tier Register-Threadgroup Memory Decomposition
by: Bergach, Mohamed Amine
Published: (2026)
by: Bergach, Mohamed Amine
Published: (2026)
gigiProfiler: Diagnosing Performance Issues by Uncovering Application Resource Bottlenecks
by: Hu, Yigong, et al.
Published: (2025)
by: Hu, Yigong, et al.
Published: (2025)
A Microbenchmark Framework for Performance Evaluation of OpenMP Target Offloading
by: Atif, Mohammad, et al.
Published: (2025)
by: Atif, Mohammad, et al.
Published: (2025)
Interpreting Performance Profiles with Deep Learning
by: Liu, Zhuoran
Published: (2025)
by: Liu, Zhuoran
Published: (2025)
Understanding the Performance Horizon of the Latest ML Workloads with NonGEMM Workloads
by: Karami, Rachid, et al.
Published: (2024)
by: Karami, Rachid, et al.
Published: (2024)
Ecoscape: Fault Tolerance Benchmark for Adaptive Remediation Strategies in Real-Time Edge ML
by: Reiter, Hendrik, et al.
Published: (2025)
by: Reiter, Hendrik, et al.
Published: (2025)
Two Criteria for Performance Analysis of Optimization Algorithms
by: Jing, Yunpeng, et al.
Published: (2024)
by: Jing, Yunpeng, et al.
Published: (2024)
DeepContext: A Context-aware, Cross-platform, and Cross-framework Tool for Performance Profiling and Analysis of Deep Learning Workloads
by: Zhao, Qidong, et al.
Published: (2024)
by: Zhao, Qidong, et al.
Published: (2024)
Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML Fusion
by: Nasr-Esfahany, Arash, et al.
Published: (2025)
by: Nasr-Esfahany, Arash, et al.
Published: (2025)
Static Reuse Profile Estimation for Array Applications
by: Razzak, Abdur, et al.
Published: (2024)
by: Razzak, Abdur, et al.
Published: (2024)
Forecasting GPU Performance for Deep Learning Training and Inference
by: Lee, Seonho, et al.
Published: (2024)
by: Lee, Seonho, et al.
Published: (2024)
Rethinking Temporal Models for TinyML: LSTM versus 1D-CNN in Resource-Constrained Devices
by: Saha, Bidyut, et al.
Published: (2026)
by: Saha, Bidyut, et al.
Published: (2026)
A Controlled Study of Memory Hierarchy Transitions in Quantum Circuit Simulation on Apple M4 Pro Unified Memory Architecture
by: Pratipat, Gyan
Published: (2026)
by: Pratipat, Gyan
Published: (2026)
Silicon Showdown: Performance, Efficiency, and Ecosystem Barriers in Consumer-Grade LLM Inference
by: Javat, Abdurrahman, et al.
Published: (2026)
by: Javat, Abdurrahman, et al.
Published: (2026)
Static Estimation of Reuse Profiles for Arrays in Nested Loops
by: Razzak, Abdur, et al.
Published: (2025)
by: Razzak, Abdur, et al.
Published: (2025)
Enabling Dynamic Sparsity in Quantized LLM Inference
by: Wang, Rongxiang, et al.
Published: (2025)
by: Wang, Rongxiang, et al.
Published: (2025)
PerfDojo: Automated ML Library Generation for Heterogeneous Architectures
by: Ivanov, Andrei, et al.
Published: (2025)
by: Ivanov, Andrei, et al.
Published: (2025)
Evaluating the Performance of the DeepSeek Model in Confidential Computing Environment
by: Dong, Ben, et al.
Published: (2025)
by: Dong, Ben, et al.
Published: (2025)
Atys: An Efficient Profiling Framework for Identifying Hotspot Functions in Large-scale Cloud Microservices
by: Sun, Jiaqi, et al.
Published: (2025)
by: Sun, Jiaqi, et al.
Published: (2025)
Systematic Performance Evaluation Framework for LEO Mega-Constellation Satellite Networks
by: Wang, Yu, et al.
Published: (2024)
by: Wang, Yu, et al.
Published: (2024)
Similar Items
-
Profiling Large Language Model Inference on Apple Silicon: A Quantization Perspective
by: Benazir, Afsara, et al.
Published: (2025) -
Accelerating Sparse Ternary GEMM for Quantized ML on Apple Silicon
by: Lipshitz, Baraq, et al.
Published: (2025) -
RWKV-edge: Deeply Compressed RWKV for Resource-Constrained Devices
by: Choe, Wonkyo, et al.
Published: (2024) -
WhisperFlow: speech foundation models in real time
by: Wang, Rongxiang, et al.
Published: (2024) -
Evaluation of Domain-Specific Architectures for General-Purpose Applications in Apple Silicon
by: López, Álvaro Corrochano, et al.
Published: (2025)