EDAN: Towards Understanding Memory Parallelism and Latency Sensitivity in HPC
Fuente:
arXiv
Saved in:
| Main Authors: | Shen, Siyuan, Khalilov, Mikhail, Gianinazzi, Lukas, Schneider, Timo, Chrapek, Marcin, Dayal, Jai, Gajbe, Manisha, Wisniewski, Robert, Hoefler, Torsten |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLAMP: Assessing Network Latency Tolerance of HPC Applications with Linear Programming
by: Shen, Siyuan, et al.
Published: (2024)
by: Shen, Siyuan, et al.
Published: (2024)
PerfDojo: Automated ML Library Generation for Heterogeneous Architectures
by: Ivanov, Andrei, et al.
Published: (2025)
by: Ivanov, Andrei, et al.
Published: (2025)
Confidential LLM Inference: Performance and Cost Across CPU and GPU TEEs
by: Chrapek, Marcin, et al.
Published: (2025)
by: Chrapek, Marcin, et al.
Published: (2025)
FPsPIN: An FPGA-based Open-Hardware Research Platform for Processing in the Network
by: Schneider, Timo, et al.
Published: (2024)
by: Schneider, Timo, et al.
Published: (2024)
Understanding Data Movement in Tightly Coupled Heterogeneous Systems: A Case Study with the Grace Hopper Superchip
by: Fusco, Luigi, et al.
Published: (2024)
by: Fusco, Luigi, et al.
Published: (2024)
Inductive Loop Analysis for Practical HPC Application Optimization
by: Schaad, Philipp, et al.
Published: (2025)
by: Schaad, Philipp, et al.
Published: (2025)
Near-Optimal Wafer-Scale Reduce
by: Luczynski, Piotr, et al.
Published: (2024)
by: Luczynski, Piotr, et al.
Published: (2024)
A Priori Loop Nest Normalization: Automatic Loop Scheduling in Complex Applications
by: Trümper, Lukas, et al.
Published: (2024)
by: Trümper, Lukas, et al.
Published: (2024)
Software Resource Disaggregation for HPC with Serverless Computing
by: Copik, Marcin, et al.
Published: (2024)
by: Copik, Marcin, et al.
Published: (2024)
DaCe AD: Unifying High-Performance Automatic Differentiation for Machine Learning and Scientific Computing
by: Boudaoud, Afif, et al.
Published: (2025)
by: Boudaoud, Afif, et al.
Published: (2025)
Hazel: Secure and Efficient Disaggregated Storage
by: Chrapek, Marcin, et al.
Published: (2025)
by: Chrapek, Marcin, et al.
Published: (2025)
Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AI
by: Khalilov, Mikhail, et al.
Published: (2024)
by: Khalilov, Mikhail, et al.
Published: (2024)
OSMOSIS: Enabling Multi-Tenancy in Datacenter SmartNICs
by: Khalilov, Mikhail, et al.
Published: (2023)
by: Khalilov, Mikhail, et al.
Published: (2023)
Iterating Pointers: Enabling Static Analysis for Loop-based Pointers
by: Lepori, Andrea, et al.
Published: (2025)
by: Lepori, Andrea, et al.
Published: (2025)
Long-term Monitoring of Kernel and Hardware Events to Understand Latency Variance
by: Zhou, Fang, et al.
Published: (2026)
by: Zhou, Fang, et al.
Published: (2026)
In-Network Collective Operations: Game Changer or Challenge for AI Workloads?
by: Hoefler, Torsten, et al.
Published: (2026)
by: Hoefler, Torsten, et al.
Published: (2026)
Analysis and Evaluation of Using Microsecond-Latency Memory for In-Memory Indices and Caches in SSD-Based Key-Value Stores
by: Bando, Yosuke, et al.
Published: (2025)
by: Bando, Yosuke, et al.
Published: (2025)
LLload: An Easy-to-Use HPC Utilization Tool
by: Byun, Chansup, et al.
Published: (2024)
by: Byun, Chansup, et al.
Published: (2024)
Streaming Data in HPC Workflows Using ADIOS
by: Eisenhauer, Greg, et al.
Published: (2024)
by: Eisenhauer, Greg, et al.
Published: (2024)
SpaDA: A Spatial Dataflow Architecture Programming Language
by: Gianinazzi, Lukas, et al.
Published: (2025)
by: Gianinazzi, Lukas, et al.
Published: (2025)
Parallel I/O Characterization and Optimization on Large-Scale HPC Systems: A 360-Degree Survey
by: Ather, Hammad, et al.
Published: (2024)
by: Ather, Hammad, et al.
Published: (2024)
ADELIA: Automatic Differentiation for Efficient Laplace Inference Approximations
by: Boudaoud, Afif, et al.
Published: (2026)
by: Boudaoud, Afif, et al.
Published: (2026)
COMPASS: A Unified Decision-Intelligence System for Navigating Performance Trade-off in HPC
by: Lahiry, Ankur, et al.
Published: (2026)
by: Lahiry, Ankur, et al.
Published: (2026)
WANDER: An Explainable Decision-Support Framework for HPC
by: Lahiry, Ankur, et al.
Published: (2025)
by: Lahiry, Ankur, et al.
Published: (2025)
HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing
by: Huang, Haochen, et al.
Published: (2025)
by: Huang, Haochen, et al.
Published: (2025)
PICO: Performance Insights for Collective Operations
by: Pasqualoni, Saverio, et al.
Published: (2025)
by: Pasqualoni, Saverio, et al.
Published: (2025)
How Much Parallelism Is "Free"? A Principle of Near-Free Parallelism for Parallel Decoding
by: He, Minghua, et al.
Published: (2026)
by: He, Minghua, et al.
Published: (2026)
Latency and Privacy-Aware Resource Allocation in Vehicular Edge Computing
by: Ahmadvand, Hossein, et al.
Published: (2025)
by: Ahmadvand, Hossein, et al.
Published: (2025)
Energy Concerns with HPC Systems and Applications
by: Nana, Roblex, et al.
Published: (2023)
by: Nana, Roblex, et al.
Published: (2023)
Usability Evaluation of Cloud for HPC Applications
by: Sochat, Vanessa, et al.
Published: (2025)
by: Sochat, Vanessa, et al.
Published: (2025)
MLKAPS: Machine Learning and Adaptive Sampling for HPC Kernel Auto-tuning
by: Jam, Mathys, et al.
Published: (2025)
by: Jam, Mathys, et al.
Published: (2025)
Heuristic-Based Merging of HPC Traces to Extend Hardware Counter Coverage
by: Aubach, Júlia Orteu, et al.
Published: (2026)
by: Aubach, Júlia Orteu, et al.
Published: (2026)
A Latency-Constrained, Gated Recurrent Unit (GRU) Implementation in the Versal AI Engine
by: Sapkas, M., et al.
Published: (2025)
by: Sapkas, M., et al.
Published: (2025)
PM2Lat: Highly Accurate and Generalized Prediction of DNN Execution Latency on GPUs
by: Le, Truong-Thanh, et al.
Published: (2026)
by: Le, Truong-Thanh, et al.
Published: (2026)
Cache Blocking of Distributed-Memory Parallel Matrix Power Kernels
by: Lacey, Dane C., et al.
Published: (2024)
by: Lacey, Dane C., et al.
Published: (2024)
Denoising Application Performance Models with Noise-Resilient Priors
by: de Morais, Gustavo, et al.
Published: (2025)
by: de Morais, Gustavo, et al.
Published: (2025)
Modeling and Controlling Many-Core HPC Processors: an Alternative to PID and Moving Average Algorithms
by: Bambini, Giovanni, et al.
Published: (2024)
by: Bambini, Giovanni, et al.
Published: (2024)
Motion-to-Motion Latency Measurement Framework for Connected and Autonomous Vehicle Teleoperation
by: Provost, François, et al.
Published: (2025)
by: Provost, François, et al.
Published: (2025)
An Experimental Study of Low-Latency Video Streaming over 5G
by: Khan, Imran, et al.
Published: (2024)
by: Khan, Imran, et al.
Published: (2024)
Extrae.jl: Julia bindings for the Extrae HPC Profiler
by: Sanchez-Ramirez, Sergio, et al.
Published: (2025)
by: Sanchez-Ramirez, Sergio, et al.
Published: (2025)
Similar Items
-
LLAMP: Assessing Network Latency Tolerance of HPC Applications with Linear Programming
by: Shen, Siyuan, et al.
Published: (2024) -
PerfDojo: Automated ML Library Generation for Heterogeneous Architectures
by: Ivanov, Andrei, et al.
Published: (2025) -
Confidential LLM Inference: Performance and Cost Across CPU and GPU TEEs
by: Chrapek, Marcin, et al.
Published: (2025) -
FPsPIN: An FPGA-based Open-Hardware Research Platform for Processing in the Network
by: Schneider, Timo, et al.
Published: (2024) -
Understanding Data Movement in Tightly Coupled Heterogeneous Systems: A Case Study with the Grace Hopper Superchip
by: Fusco, Luigi, et al.
Published: (2024)