Towards An Approach to Identify Divergences in Hardware Designs for HPC Workloads
Fuente:
arXiv
Saved in:
| Main Authors: | Popovici, Doru Thom, Vega, Mario, Ioannou, Angelos, Chaix, Fabien, Mosuli, Dania, Reasoner, Blair, Nguyen, Tan, Yang, Xiaokun, Shalf, John |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Educating for Hardware Specialization in the Chiplet Era: A Path for the HPC Community
by: Yoshii, Kazutomo, et al.
Published: (2024)
by: Yoshii, Kazutomo, et al.
Published: (2024)
Towards Efficient Neuro-Symbolic AI: From Workload Characterization to Hardware Architecture
by: Wan, Zishen, et al.
Published: (2024)
by: Wan, Zishen, et al.
Published: (2024)
Multi-Objective Hardware-Mapping Co-Optimisation for Multi-DNN Workloads on Chiplet-based Accelerators
by: Das, Abhijit, et al.
Published: (2022)
by: Das, Abhijit, et al.
Published: (2022)
Towards Efficient IMC Accelerator Design Through Joint Hardware-Workload Co-optimization
by: Krestinskaya, Olga, et al.
Published: (2024)
by: Krestinskaya, Olga, et al.
Published: (2024)
Toward Open-Source Chiplets for HPC and AI: Occamy and Beyond
by: Scheffler, Paul, et al.
Published: (2025)
by: Scheffler, Paul, et al.
Published: (2025)
ControlPULP: A RISC-V On-Chip Parallel Power Controller for Many-Core HPC Processors with FPGA-Based Hardware-In-The-Loop Power and Thermal Emulation
by: Ottaviano, Alessandro, et al.
Published: (2023)
by: Ottaviano, Alessandro, et al.
Published: (2023)
Workload Characterization for Branch Predictability
by: Vikas, FNU, et al.
Published: (2025)
by: Vikas, FNU, et al.
Published: (2025)
SLTarch: Towards Scalable Point-Based Neural Rendering by Taming Workload Imbalance and Memory Irregularity
by: Li, Xingyang, et al.
Published: (2025)
by: Li, Xingyang, et al.
Published: (2025)
Towards Efficient LUT-based PIM: A Scalable and Low-Power Approach for Modern Workloads
by: Khabbazan, Bahareh, et al.
Published: (2025)
by: Khabbazan, Bahareh, et al.
Published: (2025)
Cocco: Hardware-Mapping Co-Exploration towards Memory Capacity-Communication Optimization
by: Tan, Zhanhong, et al.
Published: (2024)
by: Tan, Zhanhong, et al.
Published: (2024)
Towards the Certification of Hybrid Architectures: Analysing Interference on Hardware Accelerators through PML
by: Lesage, Benjamin, et al.
Published: (2024)
by: Lesage, Benjamin, et al.
Published: (2024)
Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis
by: Chu, Xiaoyu, et al.
Published: (2024)
by: Chu, Xiaoyu, et al.
Published: (2024)
Affordable HPC: Leveraging Small Clusters for Big Data and Graph Computing
by: Wu, Ruilong, et al.
Published: (2024)
by: Wu, Ruilong, et al.
Published: (2024)
Energy-Efficient FPGA Framework for Non-Quantized Convolutional Neural Networks
by: Athanasiadis, Angelos, et al.
Published: (2025)
by: Athanasiadis, Angelos, et al.
Published: (2025)
A Survey on Deep Learning Hardware Accelerators for Heterogeneous HPC Platforms
by: Silvano, Cristina, et al.
Published: (2023)
by: Silvano, Cristina, et al.
Published: (2023)
Architectural Classification of XR Workloads: Cross-Layer Archetypes and Implications
by: Shi, Xinyu, et al.
Published: (2026)
by: Shi, Xinyu, et al.
Published: (2026)
Allspark: Workload Orchestration for Visual Transformers on Processing In-Memory Systems
by: Ge, Mengke, et al.
Published: (2024)
by: Ge, Mengke, et al.
Published: (2024)
Towards LLM-based Root Cause Analysis of Hardware Design Failures
by: Qiu, Siyu, et al.
Published: (2025)
by: Qiu, Siyu, et al.
Published: (2025)
Nexus Machine: An Active Message Inspired Reconfigurable Architecture for Irregular Workloads
by: Juneja, Rohan, et al.
Published: (2025)
by: Juneja, Rohan, et al.
Published: (2025)
Communication Characterization of AI Workloads for Large-scale Multi-chiplet Accelerators
by: Musavi, Mariam, et al.
Published: (2024)
by: Musavi, Mariam, et al.
Published: (2024)
TROOP: At-the-Roofline Performance for Vector Processors on Low Operational Intensity Workloads
by: Purayil, Navaneeth Kunhi, et al.
Published: (2025)
by: Purayil, Navaneeth Kunhi, et al.
Published: (2025)
Calibrating DRAMPower Model for HPC: A Runtime Perspective from Real-Time Measurements
by: Shi, Xinyu, et al.
Published: (2024)
by: Shi, Xinyu, et al.
Published: (2024)
Sustainable Hardware Specialization
by: Dangi, Pranav, et al.
Published: (2024)
by: Dangi, Pranav, et al.
Published: (2024)
Workload-Aware Hardware Accelerator Mining for Distributed Deep Learning Training
by: Adnan, Muhammad, et al.
Published: (2024)
by: Adnan, Muhammad, et al.
Published: (2024)
EnergAIzer: Fast and Accurate GPU Power Estimation Framework for AI Workloads
by: Lee, Kyungmi, et al.
Published: (2026)
by: Lee, Kyungmi, et al.
Published: (2026)
Messaging-based Adaptive Vector Computing (MAVeC) Accelerator for AI Workloads
by: Chowdhury, Md. Rownak Hossain, et al.
Published: (2024)
by: Chowdhury, Md. Rownak Hossain, et al.
Published: (2024)
elasticAI.explorer: Towards a Unified End-to-End Framework for Hardware-Aware Neural Architecture Search
by: Maman, Natalie, et al.
Published: (2026)
by: Maman, Natalie, et al.
Published: (2026)
Characterization of Real Communication Patterns and Congestion Dynamics in HPC Interconnection Networks
by: de La Rosa, Miguel Sánchez, et al.
Published: (2026)
by: de La Rosa, Miguel Sánchez, et al.
Published: (2026)
CIMinus: Empowering Sparse DNN Workloads Modeling and Exploration on SRAM-based CIM Architectures
by: Qi, Yingjie, et al.
Published: (2025)
by: Qi, Yingjie, et al.
Published: (2025)
Mapping Space Exploration for Multi-Chiplet Accelerators Targeting LLM Inference Serving Workloads
by: Li, Boyu, et al.
Published: (2025)
by: Li, Boyu, et al.
Published: (2025)
Apple vs. Oranges: Evaluating the Apple Silicon M-Series SoCs for HPC Performance and Efficiency
by: Hübner, Paul, et al.
Published: (2025)
by: Hübner, Paul, et al.
Published: (2025)
MERE: Hardware-Software Co-Design for Masking Cache Miss Latency in Embedded Processors
by: You, Dean, et al.
Published: (2025)
by: You, Dean, et al.
Published: (2025)
CoroAMU: Unleashing Memory-Driven Coroutines through Latency-Aware Decoupled Operations
by: Jiang, Zhuolun, et al.
Published: (2025)
by: Jiang, Zhuolun, et al.
Published: (2025)
Workload-Aware Early-Stage Power Delivery Network Optimization via Architectural Power Traces
by: Hayes, Oran, et al.
Published: (2026)
by: Hayes, Oran, et al.
Published: (2026)
3D-TrIM: A Memory-Efficient Spatial Computing Architecture for Convolution Workloads
by: Sestito, Cristian, et al.
Published: (2025)
by: Sestito, Cristian, et al.
Published: (2025)
HCiM: ADC-Less Hybrid Analog-Digital Compute in Memory Accelerator for Deep Learning Workloads
by: Negi, Shubham, et al.
Published: (2024)
by: Negi, Shubham, et al.
Published: (2024)
Spatzformer: An Efficient Reconfigurable Dual-Core RISC-V V Cluster for Mixed Scalar-Vector Workloads
by: Perotti, Matteo, et al.
Published: (2024)
by: Perotti, Matteo, et al.
Published: (2024)
THERMOS: Thermally-Aware Multi-Objective Scheduling of AI Workloads on Heterogeneous Multi-Chiplet PIM Architectures
by: Kanani, Alish, et al.
Published: (2025)
by: Kanani, Alish, et al.
Published: (2025)
Analyzing and Improving Hardware Modeling of Accel-Sim
by: Huerta, Rodrigo, et al.
Published: (2024)
by: Huerta, Rodrigo, et al.
Published: (2024)
Hardware and software build flow with SoCMake
by: Pejašinović, Risto, et al.
Published: (2025)
by: Pejašinović, Risto, et al.
Published: (2025)
Similar Items
-
Educating for Hardware Specialization in the Chiplet Era: A Path for the HPC Community
by: Yoshii, Kazutomo, et al.
Published: (2024) -
Towards Efficient Neuro-Symbolic AI: From Workload Characterization to Hardware Architecture
by: Wan, Zishen, et al.
Published: (2024) -
Multi-Objective Hardware-Mapping Co-Optimisation for Multi-DNN Workloads on Chiplet-based Accelerators
by: Das, Abhijit, et al.
Published: (2022) -
Towards Efficient IMC Accelerator Design Through Joint Hardware-Workload Co-optimization
by: Krestinskaya, Olga, et al.
Published: (2024) -
Toward Open-Source Chiplets for HPC and AI: Occamy and Beyond
by: Scheffler, Paul, et al.
Published: (2025)