Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML Fusion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nasr-Esfahany, Arash, Alizadeh, Mohammad, Lee, Victor, Alam, Hanna, Coon, Brett W., Culler, David, Dadu, Vidushi, Dixon, Martin, Levy, Henry M., Pandey, Santosh, Ranganathan, Parthasarathy, Yazdanbakhsh, Amir |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Tao: Re-Thinking DL-based Microarchitecture Simulation
von: Pandey, Santosh, et al.
Veröffentlicht: (2024)
von: Pandey, Santosh, et al.
Veröffentlicht: (2024)
Beyond Moore's Law: Harnessing the Redshift of Generative AI with Effective Hardware-Software Co-Design
von: Yazdanbakhsh, Amir
Veröffentlicht: (2025)
von: Yazdanbakhsh, Amir
Veröffentlicht: (2025)
Silent Data Corruption by 10x Test Escapes Threatens Reliable Computing
von: Mitra, Subhasish, et al.
Veröffentlicht: (2025)
von: Mitra, Subhasish, et al.
Veröffentlicht: (2025)
Anatomy of the gem5 Simulator: AtomicSimpleCPU, TimingSimpleCPU, O3CPU, and Their Interaction with the Ruby Memory System
von: Söderström, Johan, et al.
Veröffentlicht: (2025)
von: Söderström, Johan, et al.
Veröffentlicht: (2025)
SuperUROP: An FPGA-Based Spatial Accelerator for Sparse Matrix Operations
von: Parthasarathy, Rishab
Veröffentlicht: (2025)
von: Parthasarathy, Rishab
Veröffentlicht: (2025)
DaCapo: Accelerating Continuous Learning in Autonomous Systems for Video Analytics
von: Kim, Yoonsung, et al.
Veröffentlicht: (2024)
von: Kim, Yoonsung, et al.
Veröffentlicht: (2024)
Life-Cycle Emissions of AI Hardware: A Cradle-To-Grave Approach and Generational Trends
von: Schneider, Ian, et al.
Veröffentlicht: (2025)
von: Schneider, Ian, et al.
Veröffentlicht: (2025)
NeuroBlend: Towards Low-Power yet Accurate Neural Network-Based Inference Engine Blending Binary and Fixed-Point Convolutions
von: Fayyazi, Arash, et al.
Veröffentlicht: (2023)
von: Fayyazi, Arash, et al.
Veröffentlicht: (2023)
Further Evaluations of a Didactic CPU Visual Simulator (CPUVSIM)
von: Cortinovis, Renato, et al.
Veröffentlicht: (2024)
von: Cortinovis, Renato, et al.
Veröffentlicht: (2024)
Branch Prediction in Hardcaml for a RISC-V 32im CPU
von: Saveau, Alex
Veröffentlicht: (2023)
von: Saveau, Alex
Veröffentlicht: (2023)
A$^3$PIM: An Automated, Analytic and Accurate Processing-in-Memory Offloader
von: Jiang, Qingcai, et al.
Veröffentlicht: (2024)
von: Jiang, Qingcai, et al.
Veröffentlicht: (2024)
always_comm: An FPGA-based Hardware Accelerator for Audio/Video Compression and Transmission
von: Parthasarathy, Rishab, et al.
Veröffentlicht: (2025)
von: Parthasarathy, Rishab, et al.
Veröffentlicht: (2025)
SPEC CPU2026: Characterization, Representativeness, and Cross-Suite Comparison
von: Li, Ruihao, et al.
Veröffentlicht: (2026)
von: Li, Ruihao, et al.
Veröffentlicht: (2026)
Extending CPU-less parallel execution of lambda calculus in digital logic with lists and arithmetic
von: Fitchett, Harry, et al.
Veröffentlicht: (2026)
von: Fitchett, Harry, et al.
Veröffentlicht: (2026)
Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers
von: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Veröffentlicht: (2024)
von: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Veröffentlicht: (2024)
ISAAC: Intelligent, Scalable, Agile, and Accelerated CPU Verification via LLM-aided FPGA Parallelism
von: Sun, Jialin, et al.
Veröffentlicht: (2025)
von: Sun, Jialin, et al.
Veröffentlicht: (2025)
QiMeng-CPU-v2: Automated Superscalar Processor Design by Learning Data Dependencies
von: Cheng, Shuyao, et al.
Veröffentlicht: (2025)
von: Cheng, Shuyao, et al.
Veröffentlicht: (2025)
CPU Simulation with Ranked Set Sampling and Repeated Subsampling
von: Ekman, Magnus
Veröffentlicht: (2026)
von: Ekman, Magnus
Veröffentlicht: (2026)
CPU Simulation Using Two-Phase Stratified Sampling
von: Ekman, Magnus
Veröffentlicht: (2026)
von: Ekman, Magnus
Veröffentlicht: (2026)
EdgeLLM: A Highly Efficient CPU-FPGA Heterogeneous Edge Accelerator for Large Language Models
von: Huang, Mingqiang, et al.
Veröffentlicht: (2024)
von: Huang, Mingqiang, et al.
Veröffentlicht: (2024)
AgileWatts: An Energy-Efficient CPU Core Idle-State Architecture for Latency-Sensitive Server Applications
von: Yahya, Jawad Haj, et al.
Veröffentlicht: (2022)
von: Yahya, Jawad Haj, et al.
Veröffentlicht: (2022)
In-Storage Domain-Specific Acceleration for Serverless Computing
von: Mahapatra, Rohan, et al.
Veröffentlicht: (2023)
von: Mahapatra, Rohan, et al.
Veröffentlicht: (2023)
Towards CPU Performance Prediction: New Challenge Benchmark Dataset and Novel Approach
von: Liu, Xiaoman
Veröffentlicht: (2024)
von: Liu, Xiaoman
Veröffentlicht: (2024)
CPU-Based Layout Design for Picker-to-Parts Pallet Warehouses
von: Looms, Timo, et al.
Veröffentlicht: (2025)
von: Looms, Timo, et al.
Veröffentlicht: (2025)
EdgeMM: Multi-Core CPU with Heterogeneous AI-Extension and Activation-aware Weight Pruning for Multimodal LLMs at Edge
von: Bai, Kangbo, et al.
Veröffentlicht: (2025)
von: Bai, Kangbo, et al.
Veröffentlicht: (2025)
Confidential Computing on Heterogeneous CPU-GPU Systems: Survey and Future Directions
von: Wang, Qifan, et al.
Veröffentlicht: (2024)
von: Wang, Qifan, et al.
Veröffentlicht: (2024)
ArchPower: Dataset for Architecture-Level Power Modeling of Modern CPU Design
von: Zhang, Qijun, et al.
Veröffentlicht: (2025)
von: Zhang, Qijun, et al.
Veröffentlicht: (2025)
FERMI-ML: A Flexible and Resource-Efficient Memory-In-Situ SRAM Macro for TinyML acceleration
von: Lokhande, Mukul, et al.
Veröffentlicht: (2025)
von: Lokhande, Mukul, et al.
Veröffentlicht: (2025)
Memory Access Characterization of Large Language Models in CPU Environment and its Potential Impacts
von: Banasik, Spencer
Veröffentlicht: (2025)
von: Banasik, Spencer
Veröffentlicht: (2025)
RAGO: Systematic Performance Optimization for Retrieval-Augmented Generation Serving
von: Jiang, Wenqi, et al.
Veröffentlicht: (2025)
von: Jiang, Wenqi, et al.
Veröffentlicht: (2025)
SRAM Based Digital Custom Compute Engine for Improved Area Efficiency of AI Hardware
von: Dhakad, Narendra Singh, et al.
Veröffentlicht: (2026)
von: Dhakad, Narendra Singh, et al.
Veröffentlicht: (2026)
ADS-IMC: Accelerating Data Sorting with In-Memory Computation
von: Dhakad, Narendra Singh, et al.
Veröffentlicht: (2026)
von: Dhakad, Narendra Singh, et al.
Veröffentlicht: (2026)
Configurable Multi-Port Memory Architecture for High-Speed Data Communication
von: Dhakad, Narendra Singh, et al.
Veröffentlicht: (2024)
von: Dhakad, Narendra Singh, et al.
Veröffentlicht: (2024)
Accelerating CRONet on AMD Versal AIE-ML Engines
von: Mhatre, Kaustubh, et al.
Veröffentlicht: (2026)
von: Mhatre, Kaustubh, et al.
Veröffentlicht: (2026)
ML-based AIG Timing Prediction to Enhance Logic Optimization
von: Jiang, Wenjing, et al.
Veröffentlicht: (2024)
von: Jiang, Wenjing, et al.
Veröffentlicht: (2024)
Empowering Vector Architectures for ML: The CAMP Architecture for Matrix Multiplication
von: Nojehdeh, Mohammadreza Esmali, et al.
Veröffentlicht: (2025)
von: Nojehdeh, Mohammadreza Esmali, et al.
Veröffentlicht: (2025)
DPUConfig: Optimizing ML Inference in FPGAs Using Reinforcement Learning
von: Patras, Alexandros, et al.
Veröffentlicht: (2026)
von: Patras, Alexandros, et al.
Veröffentlicht: (2026)
KeyVisor -- A Lightweight ISA Extension for Protected Key Handles with CPU-enforced Usage Policies
von: Schwarz, Fabian, et al.
Veröffentlicht: (2024)
von: Schwarz, Fabian, et al.
Veröffentlicht: (2024)
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
von: Chung, Euijun, et al.
Veröffentlicht: (2026)
von: Chung, Euijun, et al.
Veröffentlicht: (2026)
Differentiable Initialization-Accelerated CPU-GPU Hybrid Combinatorial Scheduling
von: Liu, Mingju, et al.
Veröffentlicht: (2026)
von: Liu, Mingju, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Tao: Re-Thinking DL-based Microarchitecture Simulation
von: Pandey, Santosh, et al.
Veröffentlicht: (2024) -
Beyond Moore's Law: Harnessing the Redshift of Generative AI with Effective Hardware-Software Co-Design
von: Yazdanbakhsh, Amir
Veröffentlicht: (2025) -
Silent Data Corruption by 10x Test Escapes Threatens Reliable Computing
von: Mitra, Subhasish, et al.
Veröffentlicht: (2025) -
Anatomy of the gem5 Simulator: AtomicSimpleCPU, TimingSimpleCPU, O3CPU, and Their Interaction with the Ruby Memory System
von: Söderström, Johan, et al.
Veröffentlicht: (2025) -
SuperUROP: An FPGA-Based Spatial Accelerator for Sparse Matrix Operations
von: Parthasarathy, Rishab
Veröffentlicht: (2025)