THAPI: Tracing Heterogeneous APIs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bekele, Solomon, Vivas, Aurelio, Applencourt, Thomas, Muralidharan, Servesh, Allen, Bryce, Yoshiiinst, Kazutomo, Perarnau, Swann, Videau, Brice |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
An Online Probabilistic Distributed Tracing System
von: Toslali, M., et al.
Veröffentlicht: (2024)
von: Toslali, M., et al.
Veröffentlicht: (2024)
Recorder: Comprehensive Parallel I/O Tracing and Analysis
von: Wang, Chen, et al.
Veröffentlicht: (2025)
von: Wang, Chen, et al.
Veröffentlicht: (2025)
Performance Impact of Containerized METADOCK 2 on Heterogeneous Platforms
von: Banegas-Luna, Antonio Jesús, et al.
Veröffentlicht: (2025)
von: Banegas-Luna, Antonio Jesús, et al.
Veröffentlicht: (2025)
Understanding Power Consumption Metric on Heterogeneous Memory Systems
von: Proaño, Andrès Rubio, et al.
Veröffentlicht: (2024)
von: Proaño, Andrès Rubio, et al.
Veröffentlicht: (2024)
LEO: Tracing GPU Stall Root Causes via Cross-Vendor Backward Slicing
von: Xia, Yuning, et al.
Veröffentlicht: (2026)
von: Xia, Yuning, et al.
Veröffentlicht: (2026)
Characterizing Adaptive Mesh Refinement on Heterogeneous Platforms with Parthenon-VIBE
von: Poptani, Akash, et al.
Veröffentlicht: (2025)
von: Poptani, Akash, et al.
Veröffentlicht: (2025)
Node Compass: Multilevel Tracing and Debugging of Request Executions in JavaScript-Based Web-Servers
von: Kabamba, Herve Mbikayi, et al.
Veröffentlicht: (2023)
von: Kabamba, Herve Mbikayi, et al.
Veröffentlicht: (2023)
Enhancing Performance Insight at Scale: A Heterogeneous Framework for Exascale Diagnostics
von: Grbic, Dragana
Veröffentlicht: (2026)
von: Grbic, Dragana
Veröffentlicht: (2026)
Comparing the Performance of Heterogeneous Conjugate Gradient and Cholesky Solvers on Various Hardware Using SYCL
von: Thüring, Tim, et al.
Veröffentlicht: (2026)
von: Thüring, Tim, et al.
Veröffentlicht: (2026)
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
von: Zhang, Yaozheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yaozheng, et al.
Veröffentlicht: (2025)
Online GPU Energy Optimization with Switching-Aware Bandits
von: Xu, Xiongxiao, et al.
Veröffentlicht: (2024)
von: Xu, Xiongxiao, et al.
Veröffentlicht: (2024)
Parallel I/O Characterization and Optimization on Large-Scale HPC Systems: A 360-Degree Survey
von: Ather, Hammad, et al.
Veröffentlicht: (2024)
von: Ather, Hammad, et al.
Veröffentlicht: (2024)
Arkade: k-Nearest Neighbor Search With Non-Euclidean Distances using GPU Ray Tracing
von: Mandarapu, Durga, et al.
Veröffentlicht: (2023)
von: Mandarapu, Durga, et al.
Veröffentlicht: (2023)
CloverLeaf on Intel Multi-Core CPUs: A Case Study in Write-Allocate Evasion
von: Laukemann, Jan, et al.
Veröffentlicht: (2023)
von: Laukemann, Jan, et al.
Veröffentlicht: (2023)
Alya towards Exascale: Optimal OpenACC Performance of the Navier-Stokes Finite Element Assembly on GPUs
von: Owen, Herbert, et al.
Veröffentlicht: (2024)
von: Owen, Herbert, et al.
Veröffentlicht: (2024)
Constructive community race: full-density spiking neural network model drives neuromorphic computing
von: Senk, Johanna, et al.
Veröffentlicht: (2025)
von: Senk, Johanna, et al.
Veröffentlicht: (2025)
Sustaining Exascale Performance: Lessons from HPL and HPL-MxP on Aurora
von: Goto, Kazushige, et al.
Veröffentlicht: (2026)
von: Goto, Kazushige, et al.
Veröffentlicht: (2026)
Extracting Practical, Actionable Energy Insights from Supercomputer Telemetry and Logs
von: Cornelius, Melanie, et al.
Veröffentlicht: (2025)
von: Cornelius, Melanie, et al.
Veröffentlicht: (2025)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
Profiling and optimization of multi-card GPU machine learning jobs
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
Optimal Parallel Scheduling under Concave Speedup Functions
von: Li, Chengzhang, et al.
Veröffentlicht: (2025)
von: Li, Chengzhang, et al.
Veröffentlicht: (2025)
WebAssembly and Unikernels: A Comparative Study for Serverless at the Edge
von: Besozzi, Valerio, et al.
Veröffentlicht: (2025)
von: Besozzi, Valerio, et al.
Veröffentlicht: (2025)
Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
von: Liu, Shifang, et al.
Veröffentlicht: (2025)
von: Liu, Shifang, et al.
Veröffentlicht: (2025)
Resource Management Schemes for Cloud-Native Platforms with Computing Containers of Docker and Kubernetes
von: Mao, Ying, et al.
Veröffentlicht: (2020)
von: Mao, Ying, et al.
Veröffentlicht: (2020)
Staging Blocked Evaluation over Structured Sparse Matrices
von: Das, Pratyush, et al.
Veröffentlicht: (2024)
von: Das, Pratyush, et al.
Veröffentlicht: (2024)
Cloud Performance Decomposition for Long-Term Performance Engineering: A Case Study
von: Debnath, Shimul, et al.
Veröffentlicht: (2026)
von: Debnath, Shimul, et al.
Veröffentlicht: (2026)
Collaborative Processing for Multi-Tenant Inference on Memory-Constrained Edge TPUs
von: Ng, Nathan, et al.
Veröffentlicht: (2026)
von: Ng, Nathan, et al.
Veröffentlicht: (2026)
Preliminary report: Initial evaluation of StdPar implementations on AMD GPUs for HPC
von: Lin, Wei-Chen, et al.
Veröffentlicht: (2024)
von: Lin, Wei-Chen, et al.
Veröffentlicht: (2024)
Serving Chain-structured Jobs with Large Memory Footprints with Application to Large Foundation Model Serving
von: Sun, Tingyang, et al.
Veröffentlicht: (2026)
von: Sun, Tingyang, et al.
Veröffentlicht: (2026)
Reducing Tail Latencies Through Environment- and Neighbour-aware Thread Management
von: Jeffery, Andrew, et al.
Veröffentlicht: (2024)
von: Jeffery, Andrew, et al.
Veröffentlicht: (2024)
Dissecting the software-based measurement of CPU energy consumption: a comparative analysis
von: Raffin, Guillaume, et al.
Veröffentlicht: (2024)
von: Raffin, Guillaume, et al.
Veröffentlicht: (2024)
Bridding OT and PaaS in Edge-to-Cloud Continuum
von: Barrios, Carlos J, et al.
Veröffentlicht: (2025)
von: Barrios, Carlos J, et al.
Veröffentlicht: (2025)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
von: Karfakis, George, et al.
Veröffentlicht: (2025)
von: Karfakis, George, et al.
Veröffentlicht: (2025)
Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters
von: Jain, Rutwik, et al.
Veröffentlicht: (2026)
von: Jain, Rutwik, et al.
Veröffentlicht: (2026)
Hardware-Agnostic and Insightful Efficiency Metrics for Accelerated Systems: Definition and Implementation within TALP
von: Rahimi, Ghazal, et al.
Veröffentlicht: (2026)
von: Rahimi, Ghazal, et al.
Veröffentlicht: (2026)
Beyond Thread States: Diagnosing Performance Degradation with eBPF and Thread Dynamics
von: Landau, Diogo, et al.
Veröffentlicht: (2026)
von: Landau, Diogo, et al.
Veröffentlicht: (2026)
Asymptotically Optimal Scheduling of Multiple Parallelizable Job Classes
von: Berg, Benjamin, et al.
Veröffentlicht: (2024)
von: Berg, Benjamin, et al.
Veröffentlicht: (2024)
QoSFlow: Ensuring Service Quality of Distributed Workflows Using Interpretable Sensitivity Models
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
High-Performance Portable GPU Primitives for Arbitrary Types and Operators in Julia
von: Pilliat, Emmanuel
Veröffentlicht: (2026)
von: Pilliat, Emmanuel
Veröffentlicht: (2026)
Ähnliche Einträge
-
An Online Probabilistic Distributed Tracing System
von: Toslali, M., et al.
Veröffentlicht: (2024) -
Recorder: Comprehensive Parallel I/O Tracing and Analysis
von: Wang, Chen, et al.
Veröffentlicht: (2025) -
Performance Impact of Containerized METADOCK 2 on Heterogeneous Platforms
von: Banegas-Luna, Antonio Jesús, et al.
Veröffentlicht: (2025) -
Understanding Power Consumption Metric on Heterogeneous Memory Systems
von: Proaño, Andrès Rubio, et al.
Veröffentlicht: (2024) -
LEO: Tracing GPU Stall Root Causes via Cross-Vendor Backward Slicing
von: Xia, Yuning, et al.
Veröffentlicht: (2026)