PASTA: A Modular Program Analysis Tool Framework for Accelerators
Fuente:
arXiv
Salvato in:
| Autori principali: | Lin, Mao, Jeon, Hyeran, Zhou, Keren |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
di: Lin, Mao, et al.
Pubblicazione: (2026)
di: Lin, Mao, et al.
Pubblicazione: (2026)
Automated Programmatic Performance Analysis of Parallel Programs
di: Cankur, Onur, et al.
Pubblicazione: (2024)
di: Cankur, Onur, et al.
Pubblicazione: (2024)
FAILS: A Framework for Automated Collection and Analysis of LLM Service Incidents
di: Battaglini-Fischer, Sándor, et al.
Pubblicazione: (2025)
di: Battaglini-Fischer, Sándor, et al.
Pubblicazione: (2025)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
di: Zhang, Li, et al.
Pubblicazione: (2025)
di: Zhang, Li, et al.
Pubblicazione: (2025)
Synthesizing Proxy Applications for MPI Programs
di: Luo, Jiyu, et al.
Pubblicazione: (2023)
di: Luo, Jiyu, et al.
Pubblicazione: (2023)
Accelerating Gaussian beam tracing method with dynamic parallelism on graphics processing units
di: Sheng, Zhang, et al.
Pubblicazione: (2025)
di: Sheng, Zhang, et al.
Pubblicazione: (2025)
Taking GPU Programming Models to Task for Performance Portability
di: Davis, Joshua H., et al.
Pubblicazione: (2024)
di: Davis, Joshua H., et al.
Pubblicazione: (2024)
Hardware-Agnostic and Insightful Efficiency Metrics for Accelerated Systems: Definition and Implementation within TALP
di: Rahimi, Ghazal, et al.
Pubblicazione: (2026)
di: Rahimi, Ghazal, et al.
Pubblicazione: (2026)
A Comparison of the Performance of the Molecular Dynamics Simulation Package GROMACS Implemented in the SYCL and CUDA Programming Models
di: Apanasevich, L., et al.
Pubblicazione: (2024)
di: Apanasevich, L., et al.
Pubblicazione: (2024)
AcceleratedKernels.jl: Cross-Architecture Parallel Algorithms from a Unified, Transpiled Codebase
di: Nicusan, Andrei-Leonard, et al.
Pubblicazione: (2025)
di: Nicusan, Andrei-Leonard, et al.
Pubblicazione: (2025)
Orthrus: Accelerating Multi-BFT Consensus through Concurrent Partial Ordering of Transactions (Extended Version)
di: Lyu, Hanzheng, et al.
Pubblicazione: (2024)
di: Lyu, Hanzheng, et al.
Pubblicazione: (2024)
Enhancing Performance Insight at Scale: A Heterogeneous Framework for Exascale Diagnostics
di: Grbic, Dragana
Pubblicazione: (2026)
di: Grbic, Dragana
Pubblicazione: (2026)
CGSim: A Simulation Framework for Large Scale Distributed Computing Environment
di: Vatsavai, Sairam Sri, et al.
Pubblicazione: (2025)
di: Vatsavai, Sairam Sri, et al.
Pubblicazione: (2025)
A Performance Analysis of BFT Consensus for Blockchains
di: Chan, J. D., et al.
Pubblicazione: (2024)
di: Chan, J. D., et al.
Pubblicazione: (2024)
CASA: A Framework for SLO and Carbon-Aware Autoscaling and Scheduling in Serverless Cloud Computing
di: Qi, S., et al.
Pubblicazione: (2024)
di: Qi, S., et al.
Pubblicazione: (2024)
Resource Management Schemes for Cloud-Native Platforms with Computing Containers of Docker and Kubernetes
di: Mao, Ying, et al.
Pubblicazione: (2020)
di: Mao, Ying, et al.
Pubblicazione: (2020)
A Comprehensive Analysis of Process Energy Consumption on Multi-Socket Systems with GPUs
di: León-Vega, Luis G., et al.
Pubblicazione: (2024)
di: León-Vega, Luis G., et al.
Pubblicazione: (2024)
Scalable GPU Performance Variability Analysis framework
di: Lahiry, Ankur, et al.
Pubblicazione: (2025)
di: Lahiry, Ankur, et al.
Pubblicazione: (2025)
Performance Debugging through Microarchitectural Sensitivity and Causality Analysis
di: Dutilleul, Alban, et al.
Pubblicazione: (2024)
di: Dutilleul, Alban, et al.
Pubblicazione: (2024)
Inductive Loop Analysis for Practical HPC Application Optimization
di: Schaad, Philipp, et al.
Pubblicazione: (2025)
di: Schaad, Philipp, et al.
Pubblicazione: (2025)
Recorder: Comprehensive Parallel I/O Tracing and Analysis
di: Wang, Chen, et al.
Pubblicazione: (2025)
di: Wang, Chen, et al.
Pubblicazione: (2025)
Accelerating Particle-in-Cell Monte Carlo Simulations with MPI, OpenMP/OpenACC and Asynchronous Multi-GPU Programming
di: Williams, Jeremy J., et al.
Pubblicazione: (2024)
di: Williams, Jeremy J., et al.
Pubblicazione: (2024)
Matryoshka: Optimization of Dynamic Diverse Quantum Chemistry Systems via Elastic Parallelism Transformation
di: Wang, Tuowei, et al.
Pubblicazione: (2024)
di: Wang, Tuowei, et al.
Pubblicazione: (2024)
Data-Driven Analysis to Understand GPU Hardware Resource Usage of Optimizations
di: Islam, Tanzima Z., et al.
Pubblicazione: (2024)
di: Islam, Tanzima Z., et al.
Pubblicazione: (2024)
PlantD: Performance, Latency ANalysis, and Testing for Data Pipelines -- An Open Source Measurement, Testing, and Simulation Framework
di: Bogart, Christopher, et al.
Pubblicazione: (2025)
di: Bogart, Christopher, et al.
Pubblicazione: (2025)
Chopin: An Open Source R-language Tool to Support Spatial Analysis on Parallelizable Infrastructure
di: Song, Insang, et al.
Pubblicazione: (2024)
di: Song, Insang, et al.
Pubblicazione: (2024)
Bringing Auto-tuning to HIP: Analysis of Tuning Impact and Difficulty on AMD and Nvidia GPUs
di: Lurati, Milo, et al.
Pubblicazione: (2024)
di: Lurati, Milo, et al.
Pubblicazione: (2024)
Analytic Roofline Modeling and Energy Analysis of LULESH Proxy Application on Multi-Core Clusters
di: Afzal, Ayesha, et al.
Pubblicazione: (2024)
di: Afzal, Ayesha, et al.
Pubblicazione: (2024)
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
di: Wang, Yuxin, et al.
Pubblicazione: (2024)
di: Wang, Yuxin, et al.
Pubblicazione: (2024)
Preliminary report: Initial evaluation of StdPar implementations on AMD GPUs for HPC
di: Lin, Wei-Chen, et al.
Pubblicazione: (2024)
di: Lin, Wei-Chen, et al.
Pubblicazione: (2024)
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
di: Zhao, Xuanlei, et al.
Pubblicazione: (2024)
di: Zhao, Xuanlei, et al.
Pubblicazione: (2024)
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
di: Zhang, Yaozheng, et al.
Pubblicazione: (2025)
di: Zhang, Yaozheng, et al.
Pubblicazione: (2025)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
di: Wang, Yuxin, et al.
Pubblicazione: (2023)
di: Wang, Yuxin, et al.
Pubblicazione: (2023)
WebAssembly and Unikernels: A Comparative Study for Serverless at the Edge
di: Besozzi, Valerio, et al.
Pubblicazione: (2025)
di: Besozzi, Valerio, et al.
Pubblicazione: (2025)
Sampling in Cloud Benchmarking: A Critical Review and Methodological Guidelines
di: Akbari, Saman, et al.
Pubblicazione: (2025)
di: Akbari, Saman, et al.
Pubblicazione: (2025)
Ridgeline: A 2D Roofline Model for Distributed Systems
di: Checconi, Fabio, et al.
Pubblicazione: (2022)
di: Checconi, Fabio, et al.
Pubblicazione: (2022)
A dynamic parallel method for performance optimization on hybrid CPUs
di: Yu, Luo, et al.
Pubblicazione: (2024)
di: Yu, Luo, et al.
Pubblicazione: (2024)
EfiMon: A Process Analyser for Granular Power Consumption Prediction
di: León-Vega, Luis G., et al.
Pubblicazione: (2024)
di: León-Vega, Luis G., et al.
Pubblicazione: (2024)
Automated Calibration of Parallel and Distributed Computing Simulators: A Case Study
di: McDonald, Jesse, et al.
Pubblicazione: (2024)
di: McDonald, Jesse, et al.
Pubblicazione: (2024)
Universal Workers: A Vision for Eliminating Cold Starts in Serverless Computing
di: Akbari, Saman, et al.
Pubblicazione: (2025)
di: Akbari, Saman, et al.
Pubblicazione: (2025)
Documenti analoghi
-
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
di: Lin, Mao, et al.
Pubblicazione: (2026) -
Automated Programmatic Performance Analysis of Parallel Programs
di: Cankur, Onur, et al.
Pubblicazione: (2024) -
FAILS: A Framework for Automated Collection and Analysis of LLM Service Incidents
di: Battaglini-Fischer, Sándor, et al.
Pubblicazione: (2025) -
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
di: Zhang, Li, et al.
Pubblicazione: (2025) -
Synthesizing Proxy Applications for MPI Programs
di: Luo, Jiyu, et al.
Pubblicazione: (2023)