Introducing Instruction-Accurate Simulators for Performance Estimation of Autotuning Workloads
Fuente:
arXiv
Saved in:
| Main Authors: | Pelke, Rebecca, Bosbach, Nils, Reimann, Lennart M., Leupers, Rainer |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CLSA-CIM: A Cross-Layer Scheduling Approach for Computing-in-Memory Architectures
by: Pelke, Rebecca, et al.
Published: (2024)
by: Pelke, Rebecca, et al.
Published: (2024)
Automatic Microarchitecture-Aware Custom Instruction Design for RISC-V Processors
by: Rezunov, Evgenii, et al.
Published: (2025)
by: Rezunov, Evgenii, et al.
Published: (2025)
Bridging the Gap: Physical PCI Device Integration Into SystemC-TLM Virtual Platforms
by: Bosbach, Nils, et al.
Published: (2025)
by: Bosbach, Nils, et al.
Published: (2025)
QTFlow: Quantitative Timing-Sensitive Information Flow for Security-Aware Hardware Design on RTL
by: Reimann, Lennart M., et al.
Published: (2024)
by: Reimann, Lennart M., et al.
Published: (2024)
NISTT: A Non-Intrusive SystemC-TLM 2.0 Tracing Tool
by: Bosbach, Nils, et al.
Published: (2022)
by: Bosbach, Nils, et al.
Published: (2022)
Understanding the Performance Horizon of the Latest ML Workloads with NonGEMM Workloads
by: Karami, Rachid, et al.
Published: (2024)
by: Karami, Rachid, et al.
Published: (2024)
Exploiting the Lock: Leveraging MiG-V's Logic Locking for Secret-Data Extraction
by: Reimann, Lennart M., et al.
Published: (2024)
by: Reimann, Lennart M., et al.
Published: (2024)
The Impact of Logic Locking on Confidentiality: An Automated Evaluation
by: Reimann, Lennart M., et al.
Published: (2025)
by: Reimann, Lennart M., et al.
Published: (2025)
NeuroScalar: A Deep Learning Framework for Fast, Accurate, and In-the-Wild Cycle-Level Performance Prediction
by: Wadle, Shayne, et al.
Published: (2025)
by: Wadle, Shayne, et al.
Published: (2025)
Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML Fusion
by: Nasr-Esfahany, Arash, et al.
Published: (2025)
by: Nasr-Esfahany, Arash, et al.
Published: (2025)
EnergAIzer: Fast and Accurate GPU Power Estimation Framework for AI Workloads
by: Lee, Kyungmi, et al.
Published: (2026)
by: Lee, Kyungmi, et al.
Published: (2026)
Rescaling-Aware Training for Efficient Deployment of Deep Learning Models on Full-Integer Hardware
by: Mueller, Lion, et al.
Published: (2025)
by: Mueller, Lion, et al.
Published: (2025)
SIP: Autotuning GPU Native Schedules via Stochastic Instruction Perturbation
by: He, Guoliang, et al.
Published: (2024)
by: He, Guoliang, et al.
Published: (2024)
MGS: Markov Greedy Sums for Accurate Low-Bitwidth Floating-Point Accumulation
by: Natesh, Vikas, et al.
Published: (2025)
by: Natesh, Vikas, et al.
Published: (2025)
SLOFetch: Compressed-Hierarchical Instruction Prefetching for Cloud Microservices
by: Bao, Zerui, et al.
Published: (2025)
by: Bao, Zerui, et al.
Published: (2025)
ATiM: Autotuning Tensor Programs for Processing-in-DRAM
by: Shin, Yongwon, et al.
Published: (2024)
by: Shin, Yongwon, et al.
Published: (2024)
Compact Yet Highly Accurate Printed Classifiers Using Sequential Support Vector Machine Circuits
by: Sertaridis, Ilias, et al.
Published: (2025)
by: Sertaridis, Ilias, et al.
Published: (2025)
GENIE-ASI: Generative Instruction and Executable Code for Analog Subcircuit Identification
by: Pham, Phuoc, et al.
Published: (2025)
by: Pham, Phuoc, et al.
Published: (2025)
The Role of Advanced Computer Architectures in Accelerating Artificial Intelligence Workloads
by: Amin, Shahid, et al.
Published: (2025)
by: Amin, Shahid, et al.
Published: (2025)
Evaluating CUDA Tile for AI Workloads on Hopper and Blackwell GPUs
by: Yadav, Divakar Kumar, et al.
Published: (2026)
by: Yadav, Divakar Kumar, et al.
Published: (2026)
DS-CIM: Digital Stochastic Computing-In-Memory Featuring Accurate OR-Accumulation via Sample Region Remapping for Edge AI Models
by: Shao, Kunming, et al.
Published: (2026)
by: Shao, Kunming, et al.
Published: (2026)
Hardware-Aware Data and Instruction Mapping for AI Tasks: Balancing Parallelism, I/O and Memory Tradeoffs
by: Chowdhury, Md Rownak Hossain, et al.
Published: (2025)
by: Chowdhury, Md Rownak Hossain, et al.
Published: (2025)
GPU Performance Portability needs Autotuning
by: Ringlein, Burkhard, et al.
Published: (2025)
by: Ringlein, Burkhard, et al.
Published: (2025)
Tao: Re-Thinking DL-based Microarchitecture Simulation
by: Pandey, Santosh, et al.
Published: (2024)
by: Pandey, Santosh, et al.
Published: (2024)
Memory-Efficient FPGA Implementation of Stochastic Simulated Annealing
by: Shin, Duckgyu, et al.
Published: (2026)
by: Shin, Duckgyu, et al.
Published: (2026)
Accelerating Computer Architecture Simulation through Machine Learning
by: Ali, Wajid, et al.
Published: (2024)
by: Ali, Wajid, et al.
Published: (2024)
Performance Modeling and Workload Analysis of Distributed Large Language Model Training and Inference
by: Kundu, Joyjit, et al.
Published: (2024)
by: Kundu, Joyjit, et al.
Published: (2024)
Taming the Tail: NoI Topology Synthesis for Mixed DL Workloads on Chiplet-Based Accelerators
by: Shukla, Arnav, et al.
Published: (2025)
by: Shukla, Arnav, et al.
Published: (2025)
From Loop Nests to Silicon: Mapping AI Workloads onto AMD NPUs with MLIR-AIR
by: Wang, Erwei, et al.
Published: (2025)
by: Wang, Erwei, et al.
Published: (2025)
Algorithm and Hardware Co-Design for Efficient Complex-Valued Uncertainty Estimation
by: Zhang, Zehuan, et al.
Published: (2026)
by: Zhang, Zehuan, et al.
Published: (2026)
Learning Generalizable Program and Architecture Representations for Performance Modeling
by: Li, Lingda, et al.
Published: (2023)
by: Li, Lingda, et al.
Published: (2023)
Beyond Tokens: Enhancing RTL Quality Estimation via Structural Graph Learning
by: Liu, Yi, et al.
Published: (2025)
by: Liu, Yi, et al.
Published: (2025)
Improving Simulation Regression Efficiency using a Machine Learning-based Method in Design Verification
by: Gadde, Deepak Narayan, et al.
Published: (2024)
by: Gadde, Deepak Narayan, et al.
Published: (2024)
Performance Analysis of DNN Inference/Training with Convolution and non-Convolution Operations
by: Esmaeilzadeh, Hadi, et al.
Published: (2023)
by: Esmaeilzadeh, Hadi, et al.
Published: (2023)
Garibaldi: A Pairwise Instruction-Data Management for Enhancing Shared Last-Level Cache Performance in Server Workloads
by: Kwon, Jaewon, et al.
Published: (2025)
by: Kwon, Jaewon, et al.
Published: (2025)
Token-Picker: Accelerating Attention in Text Generation with Minimized Memory Transfer via Probability Estimation
by: Park, Junyoung, et al.
Published: (2024)
by: Park, Junyoung, et al.
Published: (2024)
VUSA: Virtually Upscaled Systolic Array Architecture to Exploit Unstructured Sparsity in AI Acceleration
by: Helal, Shereef, et al.
Published: (2025)
by: Helal, Shereef, et al.
Published: (2025)
CBM-Dual: A 65-nm Fully Connected Chaotic Boltzmann Machine Processor for Dual Function Simulated Annealing and Reservoir Computing
by: Yoshioka, Kanta, et al.
Published: (2026)
by: Yoshioka, Kanta, et al.
Published: (2026)
Exploring the Performance Improvement of Tensor Processing Engines through Transformation in the Bit-weight Dimension of MACs
by: Wu, Qizhe, et al.
Published: (2025)
by: Wu, Qizhe, et al.
Published: (2025)
Automatic Generation of Fast and Accurate Performance Models for Deep Neural Network Accelerators
by: Lübeck, Konstantin, et al.
Published: (2024)
by: Lübeck, Konstantin, et al.
Published: (2024)
Similar Items
-
CLSA-CIM: A Cross-Layer Scheduling Approach for Computing-in-Memory Architectures
by: Pelke, Rebecca, et al.
Published: (2024) -
Automatic Microarchitecture-Aware Custom Instruction Design for RISC-V Processors
by: Rezunov, Evgenii, et al.
Published: (2025) -
Bridging the Gap: Physical PCI Device Integration Into SystemC-TLM Virtual Platforms
by: Bosbach, Nils, et al.
Published: (2025) -
QTFlow: Quantitative Timing-Sensitive Information Flow for Security-Aware Hardware Design on RTL
by: Reimann, Lennart M., et al.
Published: (2024) -
NISTT: A Non-Intrusive SystemC-TLM 2.0 Tracing Tool
by: Bosbach, Nils, et al.
Published: (2022)