Understanding the Performance Horizon of the Latest ML Workloads with NonGEMM Workloads
Fuente:
arXiv
Guardado en:
| Autores principales: | Karami, Rachid, Kao, Sheng-Chun, Kwon, Hyoukjun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SCAR: Scheduling Multi-Model AI Workloads on Heterogeneous Multi-Chiplet Module Accelerators
por: Odema, Mohanad, et al.
Publicado: (2024)
por: Odema, Mohanad, et al.
Publicado: (2024)
D-com: Accelerating Iterative Processing to Enable Low-rank Decomposition of Activations
por: Tahmasebi, Faraz, et al.
Publicado: (2025)
por: Tahmasebi, Faraz, et al.
Publicado: (2025)
DEER: Deep Runahead for Instruction Prefetching on Modern Mobile Workloads
por: Vahdatniya, Parmida, et al.
Publicado: (2025)
por: Vahdatniya, Parmida, et al.
Publicado: (2025)
Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML Fusion
por: Nasr-Esfahany, Arash, et al.
Publicado: (2025)
por: Nasr-Esfahany, Arash, et al.
Publicado: (2025)
Characterizing and Optimizing Realistic Workloads on a Commercial Compute-in-SRAM Device
por: Zhang, Niansong, et al.
Publicado: (2025)
por: Zhang, Niansong, et al.
Publicado: (2025)
Characterizing State Space Model and Hybrid Language Model Performance with Long Context
por: Mitra, Saptarshi, et al.
Publicado: (2025)
por: Mitra, Saptarshi, et al.
Publicado: (2025)
Performance and Power: Systematic Evaluation of AI Workloads on Accelerators with CARAML
por: John, Chelsea Maria, et al.
Publicado: (2024)
por: John, Chelsea Maria, et al.
Publicado: (2024)
FlexiSAGA: A Flexible Systolic Array GEMM Accelerator for Sparse and Dense Processing
por: Müller, Mika Markus, et al.
Publicado: (2025)
por: Müller, Mika Markus, et al.
Publicado: (2025)
GreenMalloc: Allocator Optimisation for Industrial Workloads
por: Dakhama, Aidan, et al.
Publicado: (2025)
por: Dakhama, Aidan, et al.
Publicado: (2025)
Adaptive Cache Pollution Control for Large Language Model Inference Workloads Using Temporal CNN-Based Prediction and Priority-Aware Replacement
por: Liu, Songze, et al.
Publicado: (2025)
por: Liu, Songze, et al.
Publicado: (2025)
USEFUSE: Uniform Stride for Enhanced Performance in Fused Layer Architecture of Deep Neural Networks
por: Ibrahim, Muhammad Sohail, et al.
Publicado: (2024)
por: Ibrahim, Muhammad Sohail, et al.
Publicado: (2024)
Introducing Instruction-Accurate Simulators for Performance Estimation of Autotuning Workloads
por: Pelke, Rebecca, et al.
Publicado: (2025)
por: Pelke, Rebecca, et al.
Publicado: (2025)
HERMES: High-Performance RISC-V Memory Hierarchy for ML Workloads
por: Suryadevara, Pranav
Publicado: (2025)
por: Suryadevara, Pranav
Publicado: (2025)
A2Q+: Improving Accumulator-Aware Weight Quantization
por: Colbert, Ian, et al.
Publicado: (2024)
por: Colbert, Ian, et al.
Publicado: (2024)
Graph neural networks with configuration cross-attention for tensor compilers
por: Khizbullin, Dmitrii, et al.
Publicado: (2024)
por: Khizbullin, Dmitrii, et al.
Publicado: (2024)
PoTAcc: A Pipeline for End-to-End Acceleration of Power-of-Two Quantized DNNs
por: Saha, Rappy, et al.
Publicado: (2026)
por: Saha, Rappy, et al.
Publicado: (2026)
Toward A Formalized Approach for Spike Sorting Algorithms and Hardware Evaluation
por: Zhang, Tim, et al.
Publicado: (2022)
por: Zhang, Tim, et al.
Publicado: (2022)
Search Your Block Floating Point Scales!
por: Gupta, Tanmaey, et al.
Publicado: (2026)
por: Gupta, Tanmaey, et al.
Publicado: (2026)
SONIQ: System-Optimized Noise-Injected Ultra-Low-Precision Quantization with Full-Precision Parity
por: Zhou, Cyrus, et al.
Publicado: (2023)
por: Zhou, Cyrus, et al.
Publicado: (2023)
GCL-Sampler: Discovering Kernel Similarity for Sampled GPU Simulation via Graph Contrastive Learning
por: Wang, Jiaqi, et al.
Publicado: (2026)
por: Wang, Jiaqi, et al.
Publicado: (2026)
Prefill vs. Decode Bottlenecks: SRAM-Frequency Tradeoffs and the Memory-Bandwidth Ceiling
por: Atmer, Hannah, et al.
Publicado: (2025)
por: Atmer, Hannah, et al.
Publicado: (2025)
Tempus: A Temporally Scalable Resource-Invariant GEMM Streaming Framework for Versal AI Edge
por: Grailoo, M., et al.
Publicado: (2026)
por: Grailoo, M., et al.
Publicado: (2026)
Profiling Concurrent Vision Inference Workloads on NVIDIA Jetson -- Extended
por: Chakraborty, Abhinaba, et al.
Publicado: (2025)
por: Chakraborty, Abhinaba, et al.
Publicado: (2025)
Characterizing and Understanding HGNN Training on GPUs
por: Han, Dengke, et al.
Publicado: (2024)
por: Han, Dengke, et al.
Publicado: (2024)
OISMA: On-the-fly In-memory Stochastic Multiplication Architecture for Matrix-Multiplication Workloads
por: Agwa, Shady, et al.
Publicado: (2025)
por: Agwa, Shady, et al.
Publicado: (2025)
In-Network Collective Operations: Game Changer or Challenge for AI Workloads?
por: Hoefler, Torsten, et al.
Publicado: (2026)
por: Hoefler, Torsten, et al.
Publicado: (2026)
DRAGON (Differentiable Graph Execution) : A suite of Hardware Simulation and Optimization tools for Modern AI/Non-AI Workloads
por: Sethi, Khushal
Publicado: (2022)
por: Sethi, Khushal
Publicado: (2022)
Performance Implications of Multi-Chiplet Neural Processing Units on Autonomous Driving Perception
por: Odema, Mohanad, et al.
Publicado: (2024)
por: Odema, Mohanad, et al.
Publicado: (2024)
Forecasting LLM Inference Performance via Hardware-Agnostic Analytical Modeling
por: Patwari, Rajeev, et al.
Publicado: (2025)
por: Patwari, Rajeev, et al.
Publicado: (2025)
Confidential LLM Inference: Performance and Cost Across CPU and GPU TEEs
por: Chrapek, Marcin, et al.
Publicado: (2025)
por: Chrapek, Marcin, et al.
Publicado: (2025)
Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes
por: Hendria, Willy Fitra
Publicado: (2026)
por: Hendria, Willy Fitra
Publicado: (2026)
Automatic Generation of Fast and Accurate Performance Models for Deep Neural Network Accelerators
por: Lübeck, Konstantin, et al.
Publicado: (2024)
por: Lübeck, Konstantin, et al.
Publicado: (2024)
It's all about PR -- Smart Benchmarking AI Accelerators using Performance Representatives
por: Jung, Alexander Louis-Ferdinand, et al.
Publicado: (2024)
por: Jung, Alexander Louis-Ferdinand, et al.
Publicado: (2024)
DEEP-GAP: Deep-learning Evaluation of Execution Parallelism in GPU Architectural Performance
por: Palaniappan, Kathiravan
Publicado: (2026)
por: Palaniappan, Kathiravan
Publicado: (2026)
Scaling Analog Photonic Accelerators for Byte-Size, Integer General Matrix Multiply (GEMM) Kernels
por: Alo, Oluwaseun Adewunmi, et al.
Publicado: (2024)
por: Alo, Oluwaseun Adewunmi, et al.
Publicado: (2024)
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
por: Zhang, Hang, et al.
Publicado: (2025)
por: Zhang, Hang, et al.
Publicado: (2025)
A Low-Dissipation and Scalable GEMM Accelerator with Silicon Nitride Photonics
por: Karempudi, Venkata Sai Praneeth, et al.
Publicado: (2024)
por: Karempudi, Venkata Sai Praneeth, et al.
Publicado: (2024)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
por: Vellaisamy, Prabhu, et al.
Publicado: (2025)
por: Vellaisamy, Prabhu, et al.
Publicado: (2025)
NSFlow: An End-to-End FPGA Framework with Scalable Dataflow Architecture for Neuro-Symbolic AI
por: Yang, Hanchen, et al.
Publicado: (2025)
por: Yang, Hanchen, et al.
Publicado: (2025)
Design Space Exploration of Approximate Computing Techniques with a Reinforcement Learning Approach
por: Saeedi, Sepide, et al.
Publicado: (2023)
por: Saeedi, Sepide, et al.
Publicado: (2023)
Ejemplares similares
-
SCAR: Scheduling Multi-Model AI Workloads on Heterogeneous Multi-Chiplet Module Accelerators
por: Odema, Mohanad, et al.
Publicado: (2024) -
D-com: Accelerating Iterative Processing to Enable Low-rank Decomposition of Activations
por: Tahmasebi, Faraz, et al.
Publicado: (2025) -
DEER: Deep Runahead for Instruction Prefetching on Modern Mobile Workloads
por: Vahdatniya, Parmida, et al.
Publicado: (2025) -
Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML Fusion
por: Nasr-Esfahany, Arash, et al.
Publicado: (2025) -
Characterizing and Optimizing Realistic Workloads on a Commercial Compute-in-SRAM Device
por: Zhang, Niansong, et al.
Publicado: (2025)