Tao: Re-Thinking DL-based Microarchitecture Simulation
Fuente:
arXiv
Saved in:
| Main Authors: | Pandey, Santosh, Yazdanbakhsh, Amir, Liu, Hang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML Fusion
by: Nasr-Esfahany, Arash, et al.
Published: (2025)
by: Nasr-Esfahany, Arash, et al.
Published: (2025)
Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers
by: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Published: (2024)
by: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Published: (2024)
Beyond Moore's Law: Harnessing the Redshift of Generative AI with Effective Hardware-Software Co-Design
by: Yazdanbakhsh, Amir
Published: (2025)
by: Yazdanbakhsh, Amir
Published: (2025)
DaCapo: Accelerating Continuous Learning in Autonomous Systems for Video Analytics
by: Kim, Yoonsung, et al.
Published: (2024)
by: Kim, Yoonsung, et al.
Published: (2024)
Zero-Space Cost Fault Tolerance for Transformer-based Language Models on ReRAM
by: Li, Bingbing, et al.
Published: (2024)
by: Li, Bingbing, et al.
Published: (2024)
SemanticBBV: A Semantic Signature for Cross-Program Knowledge Reuse in Microarchitecture Simulation
by: Liu, Zhenguo, et al.
Published: (2025)
by: Liu, Zhenguo, et al.
Published: (2025)
Improving Simulation Regression Efficiency using a Machine Learning-based Method in Design Verification
by: Gadde, Deepak Narayan, et al.
Published: (2024)
by: Gadde, Deepak Narayan, et al.
Published: (2024)
Taming the Tail: NoI Topology Synthesis for Mixed DL Workloads on Chiplet-Based Accelerators
by: Shukla, Arnav, et al.
Published: (2025)
by: Shukla, Arnav, et al.
Published: (2025)
Accelerating Computer Architecture Simulation through Machine Learning
by: Ali, Wajid, et al.
Published: (2024)
by: Ali, Wajid, et al.
Published: (2024)
Memory-Efficient FPGA Implementation of Stochastic Simulated Annealing
by: Shin, Duckgyu, et al.
Published: (2026)
by: Shin, Duckgyu, et al.
Published: (2026)
DG-RePlAce: A Dataflow-Driven GPU-Accelerated Analytical Global Placement Framework for Machine Learning Accelerators
by: Kahng, Andrew B., et al.
Published: (2024)
by: Kahng, Andrew B., et al.
Published: (2024)
Introducing Instruction-Accurate Simulators for Performance Estimation of Autotuning Workloads
by: Pelke, Rebecca, et al.
Published: (2025)
by: Pelke, Rebecca, et al.
Published: (2025)
LLM-based AI Agent for Sizing of Analog and Mixed Signal Circuit
by: Liu, Chang, et al.
Published: (2025)
by: Liu, Chang, et al.
Published: (2025)
RLPlanner: Reinforcement Learning based Floorplanning for Chiplets with Fast Thermal Analysis
by: Duan, Yuanyuan, et al.
Published: (2023)
by: Duan, Yuanyuan, et al.
Published: (2023)
Benchmarking for Single Feature Attribution with Microarchitecture Cliffs
by: Zhen, Hao, et al.
Published: (2026)
by: Zhen, Hao, et al.
Published: (2026)
DL2Fence: Integrating Deep Learning and Frame Fusion for Enhanced Detection and Localization of Refined Denial-of-Service in Large-Scale NoCs
by: Wang, Haoyu, et al.
Published: (2024)
by: Wang, Haoyu, et al.
Published: (2024)
CBM-Dual: A 65-nm Fully Connected Chaotic Boltzmann Machine Processor for Dual Function Simulated Annealing and Reservoir Computing
by: Yoshioka, Kanta, et al.
Published: (2026)
by: Yoshioka, Kanta, et al.
Published: (2026)
QuArch: A Question-Answering Dataset for AI Agents in Computer Architecture
by: Prakash, Shvetank, et al.
Published: (2025)
by: Prakash, Shvetank, et al.
Published: (2025)
EPIM: Efficient Processing-In-Memory Accelerators based on Epitome
by: Wang, Chenyu, et al.
Published: (2023)
by: Wang, Chenyu, et al.
Published: (2023)
Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
VeriBug: An Attention-based Framework for Bug-Localization in Hardware Designs
by: Stracquadanio, Giuseppe, et al.
Published: (2024)
by: Stracquadanio, Giuseppe, et al.
Published: (2024)
LLM-USO: Large Language Model-based Universal Sizing Optimizer
by: S, Karthik Somayaji N., et al.
Published: (2025)
by: S, Karthik Somayaji N., et al.
Published: (2025)
An Analog and Digital Hybrid Attention Accelerator for Transformers with Charge-based In-memory Computing
by: Moradifirouzabadi, Ashkan, et al.
Published: (2024)
by: Moradifirouzabadi, Ashkan, et al.
Published: (2024)
DAISM: Digital Approximate In-SRAM Multiplier-based Accelerator for DNN Training and Inference
by: Sonnino, Lorenzo, et al.
Published: (2023)
by: Sonnino, Lorenzo, et al.
Published: (2023)
CHIME: Chiplet-based Heterogeneous Near-Memory Acceleration for Edge Multimodal LLM Inference
by: Chen, Yanru, et al.
Published: (2025)
by: Chen, Yanru, et al.
Published: (2025)
Toward Attention-based TinyML: A Heterogeneous Accelerated Architecture and Automated Deployment Flow
by: Wiese, Philip, et al.
Published: (2024)
by: Wiese, Philip, et al.
Published: (2024)
PASCAL: A Phase-Aware Scheduling Algorithm for Serving Reasoning-based Large Language Models
by: Cho, Eunyeong, et al.
Published: (2026)
by: Cho, Eunyeong, et al.
Published: (2026)
Bespoke Co-processor for Energy-Efficient Health Monitoring on RISC-V-based Flexible Wearables
by: Vergos, Theofanis, et al.
Published: (2025)
by: Vergos, Theofanis, et al.
Published: (2025)
PolyLUT: Learning Piecewise Polynomials for Ultra-Low Latency FPGA LUT-based Inference
by: Andronic, Marta, et al.
Published: (2023)
by: Andronic, Marta, et al.
Published: (2023)
Context-Aware Mixture-of-Experts Inference on CXL-Enabled GPU-NDP Systems
by: Fan, Zehao, et al.
Published: (2025)
by: Fan, Zehao, et al.
Published: (2025)
Efficient Message Passing Architecture for GCN Training on HBM-based FPGAs with Orthogonal Topology On-Chip Networks
by: Wu, Qizhe, et al.
Published: (2024)
by: Wu, Qizhe, et al.
Published: (2024)
Evaluating Four FPGA-accelerated Space Use Cases based on Neural Network Algorithms for On-board Inference
by: Antunes, Pedro, et al.
Published: (2026)
by: Antunes, Pedro, et al.
Published: (2026)
AutoPPA: Automated Circuit PPA Optimization via Contrastive Code-based Rule Library Learning
by: Li, Chongxiao, et al.
Published: (2026)
by: Li, Chongxiao, et al.
Published: (2026)
Self-Attention to Operator Learning-based 3D-IC Thermal Simulation
by: Huang, Zhen, et al.
Published: (2025)
by: Huang, Zhen, et al.
Published: (2025)
EEsizer: LLM-Based AI Agent for Sizing of Analog and Mixed Signal Circuit
by: Liu, Chang, et al.
Published: (2025)
by: Liu, Chang, et al.
Published: (2025)
PDA-LSTM: Knowledge-driven page data arrangement based on LSTM for LCM supression in QLC 3D NAND flash memories
by: Li, Qianhui, et al.
Published: (2025)
by: Li, Qianhui, et al.
Published: (2025)
Evaluating the Effectiveness of Microarchitectural Hardware Fault Detection for Application-Specific Requirements
by: Papadopoulos, Konstantinos-Nikolaos, et al.
Published: (2024)
by: Papadopoulos, Konstantinos-Nikolaos, et al.
Published: (2024)
Automatic Microarchitecture-Aware Custom Instruction Design for RISC-V Processors
by: Rezunov, Evgenii, et al.
Published: (2025)
by: Rezunov, Evgenii, et al.
Published: (2025)
From Characterization to Microarchitecture: Designing an Elegant and Reliable BFP-Based NPU
by: Zhang, Jie, et al.
Published: (2026)
by: Zhang, Jie, et al.
Published: (2026)
OpenACMv2: An Accuracy-Constrained Co-Optimization Framework for Approximate DCiM
by: Zhou, Yiqi, et al.
Published: (2026)
by: Zhou, Yiqi, et al.
Published: (2026)
Similar Items
-
Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML Fusion
by: Nasr-Esfahany, Arash, et al.
Published: (2025) -
Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers
by: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Published: (2024) -
Beyond Moore's Law: Harnessing the Redshift of Generative AI with Effective Hardware-Software Co-Design
by: Yazdanbakhsh, Amir
Published: (2025) -
DaCapo: Accelerating Continuous Learning in Autonomous Systems for Video Analytics
by: Kim, Yoonsung, et al.
Published: (2024) -
Zero-Space Cost Fault Tolerance for Transformer-based Language Models on ReRAM
by: Li, Bingbing, et al.
Published: (2024)