CUTEv2: Unified and Configurable Matrix Extension for Diverse CPU Architectures with Minimal Design Overhead
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ye, Jinpeng, Wang, Chongxi, Li, Wenqing, Yuan, Bin, Wang, Shiyi, Zhang, Fenglu, Yue, Junyu, Xie, Jianan, Ye, Yunhao, Deng, Haoyu, Zhou, Yingkun, Cheng, Xin, Zhang, Fuxin, Wang, Jian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EdgeMM: Multi-Core CPU with Heterogeneous AI-Extension and Activation-aware Weight Pruning for Multimodal LLMs at Edge
von: Bai, Kangbo, et al.
Veröffentlicht: (2025)
von: Bai, Kangbo, et al.
Veröffentlicht: (2025)
AgileWatts: An Energy-Efficient CPU Core Idle-State Architecture for Latency-Sensitive Server Applications
von: Yahya, Jawad Haj, et al.
Veröffentlicht: (2022)
von: Yahya, Jawad Haj, et al.
Veröffentlicht: (2022)
MX: Enhancing RISC-V's Vector ISA for Ultra-Low Overhead, Energy-Efficient Matrix Multiplication
von: Perotti, Matteo, et al.
Veröffentlicht: (2024)
von: Perotti, Matteo, et al.
Veröffentlicht: (2024)
ArchPower: Dataset for Architecture-Level Power Modeling of Modern CPU Design
von: Zhang, Qijun, et al.
Veröffentlicht: (2025)
von: Zhang, Qijun, et al.
Veröffentlicht: (2025)
RTLSeek: Boosting the LLM-Based RTL Generation with Multi-Stage Diversity-Oriented Reinforcement Learning
von: Zhang, Xinyu, et al.
Veröffentlicht: (2026)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2026)
Anatomy of the gem5 Simulator: AtomicSimpleCPU, TimingSimpleCPU, O3CPU, and Their Interaction with the Ruby Memory System
von: Söderström, Johan, et al.
Veröffentlicht: (2025)
von: Söderström, Johan, et al.
Veröffentlicht: (2025)
Overmind NSA: A Unified Neuro-Symbolic Computing Architecture with Approximate Nonlinear Activations and Preemptive Memory Bypass
von: Wang, Weilun, et al.
Veröffentlicht: (2026)
von: Wang, Weilun, et al.
Veröffentlicht: (2026)
Confidential Computing on Heterogeneous CPU-GPU Systems: Survey and Future Directions
von: Wang, Qifan, et al.
Veröffentlicht: (2024)
von: Wang, Qifan, et al.
Veröffentlicht: (2024)
RIROS: A Parallel RTL Fault SImulation FRamework with TwO-Dimensional Parallelism and Unified Schedule
von: Tang, Jiaping, et al.
Veröffentlicht: (2025)
von: Tang, Jiaping, et al.
Veröffentlicht: (2025)
Lifecycle Cost-Effectiveness Modeling for Redundancy-Enhanced Multi-Chiplet Architectures
von: Liu, Zizhen, et al.
Veröffentlicht: (2026)
von: Liu, Zizhen, et al.
Veröffentlicht: (2026)
SpArch: Efficient Architecture for Sparse Matrix Multiplication
von: Zhang, Zhekai, et al.
Veröffentlicht: (2020)
von: Zhang, Zhekai, et al.
Veröffentlicht: (2020)
Investigating Memory Failure Prediction Across CPU Architectures
von: Yu, Qiao, et al.
Veröffentlicht: (2024)
von: Yu, Qiao, et al.
Veröffentlicht: (2024)
KeyVisor -- A Lightweight ISA Extension for Protected Key Handles with CPU-enforced Usage Policies
von: Schwarz, Fabian, et al.
Veröffentlicht: (2024)
von: Schwarz, Fabian, et al.
Veröffentlicht: (2024)
D-Legion: A Scalable Many-Core Architecture for Accelerating Matrix Multiplication in Quantized LLMs
von: Abdelmaksoud, Ahmed J., et al.
Veröffentlicht: (2026)
von: Abdelmaksoud, Ahmed J., et al.
Veröffentlicht: (2026)
COBRA: Algorithm-Architecture Co-optimized Binary Transformer Accelerator for Edge Inference
von: Qiao, Ye, et al.
Veröffentlicht: (2025)
von: Qiao, Ye, et al.
Veröffentlicht: (2025)
Empowering Vector Architectures for ML: The CAMP Architecture for Matrix Multiplication
von: Nojehdeh, Mohammadreza Esmali, et al.
Veröffentlicht: (2025)
von: Nojehdeh, Mohammadreza Esmali, et al.
Veröffentlicht: (2025)
MetaDSE: A Few-shot Meta-learning Framework for Cross-workload CPU Design Space Exploration
von: Xue, Runzhen, et al.
Veröffentlicht: (2025)
von: Xue, Runzhen, et al.
Veröffentlicht: (2025)
ISAAC: Intelligent, Scalable, Agile, and Accelerated CPU Verification via LLM-aided FPGA Parallelism
von: Sun, Jialin, et al.
Veröffentlicht: (2025)
von: Sun, Jialin, et al.
Veröffentlicht: (2025)
Hermes: A Unified High-Performance NTT Architecture with Hybrid Dataflow
von: Gu, Hang, et al.
Veröffentlicht: (2026)
von: Gu, Hang, et al.
Veröffentlicht: (2026)
Reconfigurable Stream Network Architecture
von: Wang, Chengyue, et al.
Veröffentlicht: (2024)
von: Wang, Chengyue, et al.
Veröffentlicht: (2024)
ODIN-Based CPU-GPU Architecture with Replay-Driven Simulation and Emulation
von: Dorairaj, Nij, et al.
Veröffentlicht: (2026)
von: Dorairaj, Nij, et al.
Veröffentlicht: (2026)
Memory-Guided Unified Hardware Accelerator for Mixed-Precision Scientific Computing
von: Wang, Chuanzhen, et al.
Veröffentlicht: (2026)
von: Wang, Chuanzhen, et al.
Veröffentlicht: (2026)
Configurable Multi-Port Memory Architecture for High-Speed Data Communication
von: Dhakad, Narendra Singh, et al.
Veröffentlicht: (2024)
von: Dhakad, Narendra Singh, et al.
Veröffentlicht: (2024)
Think with Self-Decoupling and Self-Verification: Automated RTL Design with Backtrack-ToT
von: Chao, Zhiteng, et al.
Veröffentlicht: (2025)
von: Chao, Zhiteng, et al.
Veröffentlicht: (2025)
CellE: Automated Standard Cell Library Extension via Equality Saturation
von: Ren, Yi, et al.
Veröffentlicht: (2026)
von: Ren, Yi, et al.
Veröffentlicht: (2026)
NuRedact: Non-Uniform eFPGA Architecture for Low-Overhead and Secure IP Redaction
von: Das, Voktho, et al.
Veröffentlicht: (2026)
von: Das, Voktho, et al.
Veröffentlicht: (2026)
Energy-Oriented Computing Architecture Simulator for SNN Training
von: Ma, Yunhao, et al.
Veröffentlicht: (2025)
von: Ma, Yunhao, et al.
Veröffentlicht: (2025)
ERASER: Efficient RTL FAult Simulation Framework with Trimmed Execution Redundancy
von: Tang, Jiaping, et al.
Veröffentlicht: (2025)
von: Tang, Jiaping, et al.
Veröffentlicht: (2025)
Extend IVerilog to Support Batch RTL Fault Simulation
von: Tang, Jiaping, et al.
Veröffentlicht: (2025)
von: Tang, Jiaping, et al.
Veröffentlicht: (2025)
Pecker: Bug Localization Framework for Sequential Designs via Causal Chain Reconstruction
von: Tang, Jiaping, et al.
Veröffentlicht: (2026)
von: Tang, Jiaping, et al.
Veröffentlicht: (2026)
ARCANE: Adaptive RISC-V Cache Architecture for Near-memory Extensions
von: Petrolo, Vincenzo, et al.
Veröffentlicht: (2025)
von: Petrolo, Vincenzo, et al.
Veröffentlicht: (2025)
Further Evaluations of a Didactic CPU Visual Simulator (CPUVSIM)
von: Cortinovis, Renato, et al.
Veröffentlicht: (2024)
von: Cortinovis, Renato, et al.
Veröffentlicht: (2024)
QiMeng-CPU-v2: Automated Superscalar Processor Design by Learning Data Dependencies
von: Cheng, Shuyao, et al.
Veröffentlicht: (2025)
von: Cheng, Shuyao, et al.
Veröffentlicht: (2025)
MARCA: Mamba Accelerator with ReConfigurable Architecture
von: Li, Jinhao, et al.
Veröffentlicht: (2024)
von: Li, Jinhao, et al.
Veröffentlicht: (2024)
Branch Prediction in Hardcaml for a RISC-V 32im CPU
von: Saveau, Alex
Veröffentlicht: (2023)
von: Saveau, Alex
Veröffentlicht: (2023)
Multi-objective Optimization in CPU Design Space Exploration: Attention is All You Need
von: Xue, Runzhen, et al.
Veröffentlicht: (2024)
von: Xue, Runzhen, et al.
Veröffentlicht: (2024)
Optimized Spatial Architecture Mapping Flow for Transformer Accelerators
von: Xu, Haocheng, et al.
Veröffentlicht: (2024)
von: Xu, Haocheng, et al.
Veröffentlicht: (2024)
AutoPDR: Circuit-Aware Solver Configuration Prediction for Hardware Model Checking
von: Hu, Guangyu, et al.
Veröffentlicht: (2026)
von: Hu, Guangyu, et al.
Veröffentlicht: (2026)
ReDas: A Lightweight Architecture for Supporting Fine-Grained Reshaping and Multiple Dataflows on Systolic Array
von: Han, Meng, et al.
Veröffentlicht: (2023)
von: Han, Meng, et al.
Veröffentlicht: (2023)
SPEC CPU2026: Characterization, Representativeness, and Cross-Suite Comparison
von: Li, Ruihao, et al.
Veröffentlicht: (2026)
von: Li, Ruihao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
EdgeMM: Multi-Core CPU with Heterogeneous AI-Extension and Activation-aware Weight Pruning for Multimodal LLMs at Edge
von: Bai, Kangbo, et al.
Veröffentlicht: (2025) -
AgileWatts: An Energy-Efficient CPU Core Idle-State Architecture for Latency-Sensitive Server Applications
von: Yahya, Jawad Haj, et al.
Veröffentlicht: (2022) -
MX: Enhancing RISC-V's Vector ISA for Ultra-Low Overhead, Energy-Efficient Matrix Multiplication
von: Perotti, Matteo, et al.
Veröffentlicht: (2024) -
ArchPower: Dataset for Architecture-Level Power Modeling of Modern CPU Design
von: Zhang, Qijun, et al.
Veröffentlicht: (2025) -
RTLSeek: Boosting the LLM-Based RTL Generation with Multi-Stage Diversity-Oriented Reinforcement Learning
von: Zhang, Xinyu, et al.
Veröffentlicht: (2026)