Explainable Port Mapping Inference with Sparse Performance Counters for AMD's Zen Architectures
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ritter, Fabian, Hack, Sebastian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AMD MI300X GPU Performance Analysis
von: Ambati, Chandrish, et al.
Veröffentlicht: (2025)
von: Ambati, Chandrish, et al.
Veröffentlicht: (2025)
H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference
von: Fu, Zizhuo, et al.
Veröffentlicht: (2025)
von: Fu, Zizhuo, et al.
Veröffentlicht: (2025)
Data-Driven Power Modeling and Monitoring via Hardware Performance Counters Tracking
von: Mazzola, Sergio, et al.
Veröffentlicht: (2024)
von: Mazzola, Sergio, et al.
Veröffentlicht: (2024)
Data-Driven Power Modeling and Monitoring via Hardware Performance Counter Tracking
von: Mazzola, Sergio, et al.
Veröffentlicht: (2025)
von: Mazzola, Sergio, et al.
Veröffentlicht: (2025)
Evaluating the impact of the L3 cache size of AMD EPYC CPUs on the performance of CFD applications
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
SparseInfer: Training-free Prediction of Activation Sparsity for Fast LLM Inference
von: Shin, Jiho, et al.
Veröffentlicht: (2024)
von: Shin, Jiho, et al.
Veröffentlicht: (2024)
Dissecting Embedding Bag Performance in DLRM Inference
von: Ambati, Chandrish, et al.
Veröffentlicht: (2025)
von: Ambati, Chandrish, et al.
Veröffentlicht: (2025)
Characterize LSM-tree Compaction Performance via On-Device LLM Inference
von: Ding, Jiabiao, et al.
Veröffentlicht: (2026)
von: Ding, Jiabiao, et al.
Veröffentlicht: (2026)
Meta-Metrics and Best Practices for System-Level Inference Performance Benchmarking
von: Salaria, Shweta, et al.
Veröffentlicht: (2025)
von: Salaria, Shweta, et al.
Veröffentlicht: (2025)
MarginGate: Sparse Margin-Triggered Verification for Batch-Invariant LLM Inference
von: Chu, Kexin, et al.
Veröffentlicht: (2026)
von: Chu, Kexin, et al.
Veröffentlicht: (2026)
CounterPoint: Using Hardware Event Counters to Refute and Refine Microarchitectural Assumptions (Extended Version)
von: Lindsay, Nick, et al.
Veröffentlicht: (2026)
von: Lindsay, Nick, et al.
Veröffentlicht: (2026)
$γ$-CounterBoost: Optimizing response time tails using job type information only
von: Charlet, Nils, et al.
Veröffentlicht: (2026)
von: Charlet, Nils, et al.
Veröffentlicht: (2026)
Heuristic-Based Merging of HPC Traces to Extend Hardware Counter Coverage
von: Aubach, Júlia Orteu, et al.
Veröffentlicht: (2026)
von: Aubach, Júlia Orteu, et al.
Veröffentlicht: (2026)
Preliminary report: Initial evaluation of StdPar implementations on AMD GPUs for HPC
von: Lin, Wei-Chen, et al.
Veröffentlicht: (2024)
von: Lin, Wei-Chen, et al.
Veröffentlicht: (2024)
Forecasting GPU Performance for Deep Learning Training and Inference
von: Lee, Seonho, et al.
Veröffentlicht: (2024)
von: Lee, Seonho, et al.
Veröffentlicht: (2024)
Bringing Auto-tuning to HIP: Analysis of Tuning Impact and Difficulty on AMD and Nvidia GPUs
von: Lurati, Milo, et al.
Veröffentlicht: (2024)
von: Lurati, Milo, et al.
Veröffentlicht: (2024)
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2025)
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2025)
Seamless acceleration of Fortran intrinsics via AMD AI engines
von: Brown, Nick, et al.
Veröffentlicht: (2025)
von: Brown, Nick, et al.
Veröffentlicht: (2025)
GROMACS on AMD GPU-Based HPC Platforms: Using SYCL for Performance and Portability
von: Alekseenko, Andrey, et al.
Veröffentlicht: (2024)
von: Alekseenko, Andrey, et al.
Veröffentlicht: (2024)
Fast Entropy Decoding for Sparse MVM on GPUs
von: Schätzle, Emil, et al.
Veröffentlicht: (2026)
von: Schätzle, Emil, et al.
Veröffentlicht: (2026)
AI Application Benchmarking: Power-Aware Performance Analysis for Vision and Language Models
von: Mayr, Martin, et al.
Veröffentlicht: (2026)
von: Mayr, Martin, et al.
Veröffentlicht: (2026)
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference
von: Yao, Feiyu, et al.
Veröffentlicht: (2026)
von: Yao, Feiyu, et al.
Veröffentlicht: (2026)
An Empirical Study on How Architectural Topology Affects Microservice Performance and Energy Usage
von: Ristova, Irena, et al.
Veröffentlicht: (2026)
von: Ristova, Irena, et al.
Veröffentlicht: (2026)
Architecture Specific Generation of Large Scale Lattice Boltzmann Methods for Sparse Complex Geometries
von: Suffa, Philipp, et al.
Veröffentlicht: (2024)
von: Suffa, Philipp, et al.
Veröffentlicht: (2024)
Single 32-bit Sub-Channel DDR5 DIMMs: Architecture, Performance Bounds, and Standardisation
von: Ke, Chih-Hua
Veröffentlicht: (2026)
von: Ke, Chih-Hua
Veröffentlicht: (2026)
OSCAR-P and aMLLibrary: Profiling and Predicting the Performance of FaaS-based Applications in Computing Continua
von: Sala, Roberto, et al.
Veröffentlicht: (2024)
von: Sala, Roberto, et al.
Veröffentlicht: (2024)
MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models
von: Chitty-Venkata, Krishna Teja, et al.
Veröffentlicht: (2025)
von: Chitty-Venkata, Krishna Teja, et al.
Veröffentlicht: (2025)
Multi-DNN Inference of Sparse Models on Edge SoCs
von: Luo, Jiawei, et al.
Veröffentlicht: (2026)
von: Luo, Jiawei, et al.
Veröffentlicht: (2026)
Enhancements to P4TG: Protocols, Performance, and Automation
von: Ihle, Fabian, et al.
Veröffentlicht: (2025)
von: Ihle, Fabian, et al.
Veröffentlicht: (2025)
SparseX: Efficient Segment-Level KV Cache Sharing for Interleaved LLM Serving
von: Zhang, Quqing, et al.
Veröffentlicht: (2026)
von: Zhang, Quqing, et al.
Veröffentlicht: (2026)
Statistical Modeling and Uncertainty Estimation of LLM Inference Systems
von: Ray, Kaustabha, et al.
Veröffentlicht: (2025)
von: Ray, Kaustabha, et al.
Veröffentlicht: (2025)
GDEV-AI: A Generalized Evaluation of Deep Learning Inference Scaling and Architectural Saturation
von: Palaniappan, Kathiravan
Veröffentlicht: (2026)
von: Palaniappan, Kathiravan
Veröffentlicht: (2026)
Performance Characterization of Expert Router for Scalable LLM Inference
von: Pichlmeier, Josef, et al.
Veröffentlicht: (2024)
von: Pichlmeier, Josef, et al.
Veröffentlicht: (2024)
WANDER: An Explainable Decision-Support Framework for HPC
von: Lahiry, Ankur, et al.
Veröffentlicht: (2025)
von: Lahiry, Ankur, et al.
Veröffentlicht: (2025)
cuTeSpMM: Accelerating Sparse-Dense Matrix Multiplication using GPU Tensor Cores
von: Xiang, Lizhi, et al.
Veröffentlicht: (2025)
von: Xiang, Lizhi, et al.
Veröffentlicht: (2025)
Machine Learning-Guided Memory Optimization for DLRM Inference on Tiered Memory
von: Ren, Jie, et al.
Veröffentlicht: (2025)
von: Ren, Jie, et al.
Veröffentlicht: (2025)
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference
von: Hamadanian, Pouya, et al.
Veröffentlicht: (2025)
von: Hamadanian, Pouya, et al.
Veröffentlicht: (2025)
ADS Performance Revisited
von: Weber, Alexander, et al.
Veröffentlicht: (2024)
von: Weber, Alexander, et al.
Veröffentlicht: (2024)
Profiling Large Language Model Inference on Apple Silicon: A Quantization Perspective
von: Benazir, Afsara, et al.
Veröffentlicht: (2025)
von: Benazir, Afsara, et al.
Veröffentlicht: (2025)
Scalable Systems and Software Architectures for High-Performance Computing on cloud platforms
von: Ramesh, Risshab Srinivas
Veröffentlicht: (2024)
von: Ramesh, Risshab Srinivas
Veröffentlicht: (2024)
Ähnliche Einträge
-
AMD MI300X GPU Performance Analysis
von: Ambati, Chandrish, et al.
Veröffentlicht: (2025) -
H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference
von: Fu, Zizhuo, et al.
Veröffentlicht: (2025) -
Data-Driven Power Modeling and Monitoring via Hardware Performance Counters Tracking
von: Mazzola, Sergio, et al.
Veröffentlicht: (2024) -
Data-Driven Power Modeling and Monitoring via Hardware Performance Counter Tracking
von: Mazzola, Sergio, et al.
Veröffentlicht: (2025) -
Evaluating the impact of the L3 cache size of AMD EPYC CPUs on the performance of CFD applications
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)