Saved in:
| Main Authors: | Lee, Kyungmi, Song, Zhiye, Lee, Eun Kyung, Zhang, Xin, Eilam, Tamar, Chandrakasan, Anantha P. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.20105 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EnergyLens: Predictive Energy-Aware Exploration for Multi-GPU LLM Inference Optimization
by: Song, Zhiye, et al.
Published: (2026)
by: Song, Zhiye, et al.
Published: (2026)
Introducing Instruction-Accurate Simulators for Performance Estimation of Autotuning Workloads
by: Pelke, Rebecca, et al.
Published: (2025)
by: Pelke, Rebecca, et al.
Published: (2025)
GateKeeper-GPU: Fast and Accurate Pre-Alignment Filtering in Short Read Mapping
by: Bingöl, Zülal, et al.
Published: (2021)
by: Bingöl, Zülal, et al.
Published: (2021)
Bitwise Logic Using Phase Change Memory Devices Based on the Pinatubo Architecture
by: Aflalo, Noa, et al.
Published: (2024)
by: Aflalo, Noa, et al.
Published: (2024)
CXL-GPU: Pushing GPU Memory Boundaries with the Integration of CXL Technologies
by: Gouk, Donghyun, et al.
Published: (2025)
by: Gouk, Donghyun, et al.
Published: (2025)
Fastrack: Fast IO for Secure ML using GPU TEEs
by: Wang, Yongqin, et al.
Published: (2024)
by: Wang, Yongqin, et al.
Published: (2024)
Empirical Measurements of AI Training Power Demand on a GPU-Accelerated Node
by: Latif, Imran, et al.
Published: (2024)
by: Latif, Imran, et al.
Published: (2024)
Workload-Aware Early-Stage Power Delivery Network Optimization via Architectural Power Traces
by: Hayes, Oran, et al.
Published: (2026)
by: Hayes, Oran, et al.
Published: (2026)
Communication Characterization of AI Workloads for Large-scale Multi-chiplet Accelerators
by: Musavi, Mariam, et al.
Published: (2024)
by: Musavi, Mariam, et al.
Published: (2024)
Messaging-based Adaptive Vector Computing (MAVeC) Accelerator for AI Workloads
by: Chowdhury, Md. Rownak Hossain, et al.
Published: (2024)
by: Chowdhury, Md. Rownak Hossain, et al.
Published: (2024)
Towards Efficient LUT-based PIM: A Scalable and Low-Power Approach for Modern Workloads
by: Khabbazan, Bahareh, et al.
Published: (2025)
by: Khabbazan, Bahareh, et al.
Published: (2025)
Garibaldi: A Pairwise Instruction-Data Management for Enhancing Shared Last-Level Cache Performance in Server Workloads
by: Kwon, Jaewon, et al.
Published: (2025)
by: Kwon, Jaewon, et al.
Published: (2025)
Allspark: Workload Orchestration for Visual Transformers on Processing In-Memory Systems
by: Ge, Mengke, et al.
Published: (2024)
by: Ge, Mengke, et al.
Published: (2024)
PAI: Fast, Accurate, and Full Benchmark Performance Projection with AI
by: Johnson, Avery, et al.
Published: (2026)
by: Johnson, Avery, et al.
Published: (2026)
HH-PIM: Dynamic Optimization of Power and Performance with Heterogeneous-Hybrid PIM for Edge AI Devices
by: Jeon, Sangmin, et al.
Published: (2025)
by: Jeon, Sangmin, et al.
Published: (2025)
Workload Characterization for Branch Predictability
by: Vikas, FNU, et al.
Published: (2025)
by: Vikas, FNU, et al.
Published: (2025)
FiCABU: A Fisher-Based, Context-Adaptive Machine Unlearning Processor for Edge AI
by: Cho, Eun-Su, et al.
Published: (2025)
by: Cho, Eun-Su, et al.
Published: (2025)
THERMOS: Thermally-Aware Multi-Objective Scheduling of AI Workloads on Heterogeneous Multi-Chiplet PIM Architectures
by: Kanani, Alish, et al.
Published: (2025)
by: Kanani, Alish, et al.
Published: (2025)
CiMLoop: A Flexible, Accurate, and Fast Compute-In-Memory Modeling Tool
by: Andrulis, Tanner, et al.
Published: (2024)
by: Andrulis, Tanner, et al.
Published: (2024)
CMAX-CAMEL: A Coarse-to-Fine Adaptive, Memory-Efficient, and Low-Power Edge Processor for Contrast Maximization
by: Min, Kyeongpil, et al.
Published: (2026)
by: Min, Kyeongpil, et al.
Published: (2026)
EasyRider: Mitigating Power Transients in Datacenter-Scale Training Workloads
by: Jensen, Dillon, et al.
Published: (2026)
by: Jensen, Dillon, et al.
Published: (2026)
DCI: A Coordinated Allocation and Filling Workload-Aware Dual-Cache Allocation GNN Inference Acceleration System
by: Luo, Yi, et al.
Published: (2025)
by: Luo, Yi, et al.
Published: (2025)
Edge GPU Aware Multiple AI Model Pipeline for Accelerated MRI Reconstruction and Analysis
by: Majeed, Ashiyana Abdul, et al.
Published: (2025)
by: Majeed, Ashiyana Abdul, et al.
Published: (2025)
Towards An Approach to Identify Divergences in Hardware Designs for HPC Workloads
by: Popovici, Doru Thom, et al.
Published: (2025)
by: Popovici, Doru Thom, et al.
Published: (2025)
Architectural Classification of XR Workloads: Cross-Layer Archetypes and Implications
by: Shi, Xinyu, et al.
Published: (2026)
by: Shi, Xinyu, et al.
Published: (2026)
Empirically-Calibrated H100 Node Power Models for Reducing Uncertainty in AI Training Energy Estimation
by: Newkirk, Alex C., et al.
Published: (2025)
by: Newkirk, Alex C., et al.
Published: (2025)
Regular-Dead on Arrival: Characterizing and Protecting Against Dead-Entry TLB Misses in GPU Microarchitectures
by: Anik, Shafayat Mowla, et al.
Published: (2026)
by: Anik, Shafayat Mowla, et al.
Published: (2026)
EasyDRAM: An FPGA-based Infrastructure for Fast and Accurate End-to-End Evaluation of Emerging DRAM Techniques
by: Canpolat, Oğuzhan, et al.
Published: (2025)
by: Canpolat, Oğuzhan, et al.
Published: (2025)
PIM-malloc: A Fast and Scalable Dynamic Memory Allocator for Processing-In-Memory (PIM) Architectures
by: Lee, Dongjae, et al.
Published: (2025)
by: Lee, Dongjae, et al.
Published: (2025)
Nexus Machine: An Active Message Inspired Reconfigurable Architecture for Irregular Workloads
by: Juneja, Rohan, et al.
Published: (2025)
by: Juneja, Rohan, et al.
Published: (2025)
TROOP: At-the-Roofline Performance for Vector Processors on Low Operational Intensity Workloads
by: Purayil, Navaneeth Kunhi, et al.
Published: (2025)
by: Purayil, Navaneeth Kunhi, et al.
Published: (2025)
NeuroScalar: A Deep Learning Framework for Fast, Accurate, and In-the-Wild Cycle-Level Performance Prediction
by: Wadle, Shayne, et al.
Published: (2025)
by: Wadle, Shayne, et al.
Published: (2025)
e-GPU: An Open-Source and Configurable RISC-V Graphic Processing Unit for TinyAI Applications
by: Machetti, Simone, et al.
Published: (2025)
by: Machetti, Simone, et al.
Published: (2025)
Accelerating GenAI Workloads by Enabling RISC-V Microkernel Support in IREE
by: Ahmad, Adeel, et al.
Published: (2025)
by: Ahmad, Adeel, et al.
Published: (2025)
Towards Efficient Neuro-Symbolic AI: From Workload Characterization to Hardware Architecture
by: Wan, Zishen, et al.
Published: (2024)
by: Wan, Zishen, et al.
Published: (2024)
Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML Fusion
by: Nasr-Esfahany, Arash, et al.
Published: (2025)
by: Nasr-Esfahany, Arash, et al.
Published: (2025)
RoboGPU: Accelerating GPU Collision Detection for Robotics
by: Liu, Lufei, et al.
Published: (2026)
by: Liu, Lufei, et al.
Published: (2026)
Analyzing Modern NVIDIA GPU cores
by: Huerta, Rodrigo, et al.
Published: (2025)
by: Huerta, Rodrigo, et al.
Published: (2025)
CIMinus: Empowering Sparse DNN Workloads Modeling and Exploration on SRAM-based CIM Architectures
by: Qi, Yingjie, et al.
Published: (2025)
by: Qi, Yingjie, et al.
Published: (2025)
Mapping Space Exploration for Multi-Chiplet Accelerators Targeting LLM Inference Serving Workloads
by: Li, Boyu, et al.
Published: (2025)
by: Li, Boyu, et al.
Published: (2025)
Similar Items
-
EnergyLens: Predictive Energy-Aware Exploration for Multi-GPU LLM Inference Optimization
by: Song, Zhiye, et al.
Published: (2026) -
Introducing Instruction-Accurate Simulators for Performance Estimation of Autotuning Workloads
by: Pelke, Rebecca, et al.
Published: (2025) -
GateKeeper-GPU: Fast and Accurate Pre-Alignment Filtering in Short Read Mapping
by: Bingöl, Zülal, et al.
Published: (2021) -
Bitwise Logic Using Phase Change Memory Devices Based on the Pinatubo Architecture
by: Aflalo, Noa, et al.
Published: (2024) -
CXL-GPU: Pushing GPU Memory Boundaries with the Integration of CXL Technologies
by: Gouk, Donghyun, et al.
Published: (2025)