Saved in:
| Main Authors: | Zhao, Jerry, Grubb, Daniel, Rusch, Miles, Wei, Tianrui, Anderson, Kevin, Nikolic, Borivoje, Asanovic, Krste |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2412.00997 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NeCTAr: A Heterogeneous RISC-V SoC for Language Model Inference in Intel 16
by: Schmulbach, Viansa, et al.
Published: (2025)
by: Schmulbach, Viansa, et al.
Published: (2025)
Efficient Implementation of RISC-V Vector Permutation Instructions
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
AraOS: Analyzing the Impact of Virtual Memory Management on Vector Unit Performance
by: Perotti, Matteo, et al.
Published: (2025)
by: Perotti, Matteo, et al.
Published: (2025)
Instruction-Based Coordination of Heterogeneous Processing Units for Acceleration of DNN Inference
by: Petropoulos, Anastasios, et al.
Published: (2025)
by: Petropoulos, Anastasios, et al.
Published: (2025)
Spatz: Clustering Compact RISC-V-Based Vector Units to Maximize Computing Efficiency
by: Perotti, Matteo, et al.
Published: (2023)
by: Perotti, Matteo, et al.
Published: (2023)
BlissCam: Boosting Eye Tracking Efficiency with Learned In-Sensor Sparse Sampling
by: Feng, Yu, et al.
Published: (2024)
by: Feng, Yu, et al.
Published: (2024)
Optimising Iteration Scheduling for Full-State Vector Simulation of Quantum Circuits on FPGAs
by: Moawad, Youssef, et al.
Published: (2024)
by: Moawad, Youssef, et al.
Published: (2024)
SIP: Autotuning GPU Native Schedules via Stochastic Instruction Perturbation
by: He, Guoliang, et al.
Published: (2024)
by: He, Guoliang, et al.
Published: (2024)
FADiff: Fusion-Aware Differentiable Optimization for DNN Scheduling on Tensor Accelerators
by: Jia, Shuao, et al.
Published: (2025)
by: Jia, Shuao, et al.
Published: (2025)
IMMSched: Interruptible Multi-DNN Scheduling via Parallel Multi-Particle Optimizing Subgraph Isomorphism
by: Zhao, Boran, et al.
Published: (2026)
by: Zhao, Boran, et al.
Published: (2026)
VitaLLM: A Versatile, Ultra-Compact Ternary LLM Accelerator with Dependency-Aware Scheduling
by: Lin, Zi-Wei, et al.
Published: (2026)
by: Lin, Zi-Wei, et al.
Published: (2026)
RoSE-Opt: Robust and Efficient Analog Circuit Parameter Optimization with Knowledge-infused Reinforcement Learning
by: Cao, Weidong, et al.
Published: (2024)
by: Cao, Weidong, et al.
Published: (2024)
Jack Unit: An Area- and Energy-Efficient Multiply-Accumulate (MAC) Unit Supporting Diverse Data Formats
by: Noh, Seock-Hwan, et al.
Published: (2025)
by: Noh, Seock-Hwan, et al.
Published: (2025)
SoMa: Identifying, Exploring, and Understanding the DRAM Communication Scheduling Space for DNN Accelerators
by: Cai, Jingwei, et al.
Published: (2025)
by: Cai, Jingwei, et al.
Published: (2025)
Rethinking Compute Substrates for 3D-Stacked Near-Memory LLM Decoding: Microarchitecture-Scheduling Co-Design
by: Ai, Chenyang, et al.
Published: (2026)
by: Ai, Chenyang, et al.
Published: (2026)
Register Dispersion: Reducing the Footprint of the Vector Register File in Vector Engines of Low-Cost RISC-V CPUs
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
An Energy-Efficient Approximate Posit Multiply-Divide Unit
by: Thotli, Rishi, et al.
Published: (2026)
by: Thotli, Rishi, et al.
Published: (2026)
RTGPU: Real-Time Computing with Graphics Processing Units
by: Gheibi-Fetrat, Atiyeh, et al.
Published: (2025)
by: Gheibi-Fetrat, Atiyeh, et al.
Published: (2025)
SATA: Sparsity-Aware Scheduling for Selective Token Attention
by: Fan, Zhenkun, et al.
Published: (2026)
by: Fan, Zhenkun, et al.
Published: (2026)
Hardwired-Neurons Language Processing Units as General-Purpose Cognitive Substrates
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
RPU -- A Reasoning Processing Unit
by: Adiletta, Matthew, et al.
Published: (2026)
by: Adiletta, Matthew, et al.
Published: (2026)
NeuroAI Temporal Neural Networks (NeuTNNs): Microarchitecture and Design Framework for Specialized Neuromorphic Processing Units
by: Venkatachalam, Shanmuga, et al.
Published: (2026)
by: Venkatachalam, Shanmuga, et al.
Published: (2026)
ReGate: Enabling Power Gating in Neural Processing Units
by: Xue, Yuqi, et al.
Published: (2025)
by: Xue, Yuqi, et al.
Published: (2025)
MTU: The Multifunction Tree Unit for Accelerating Zero-Knowledge Proofs
by: Mo, Jianqiao, et al.
Published: (2025)
by: Mo, Jianqiao, et al.
Published: (2025)
Energy-Efficient QoS-Aware Scheduling for S-NUCA Many-Cores
by: Wasala, Sudam M., et al.
Published: (2025)
by: Wasala, Sudam M., et al.
Published: (2025)
Cross-Layer Design of Vector-Symbolic Computing: Bridging Cognition and Brain-Inspired Hardware Acceleration
by: Du, Shuting, et al.
Published: (2025)
by: Du, Shuting, et al.
Published: (2025)
A "New Ara" for Vector Computing: An Open Source Highly Efficient RISC-V V 1.0 Vector Processor Design
by: Perotti, Matteo, et al.
Published: (2022)
by: Perotti, Matteo, et al.
Published: (2022)
AraXL: A Physically Scalable, Ultra-Wide RISC-V Vector Processor Design for Fast and Efficient Computation on Long Vectors
by: Purayil, Navaneeth Kunhi, et al.
Published: (2025)
by: Purayil, Navaneeth Kunhi, et al.
Published: (2025)
Support Vector Machines Classification on Bendable RISC-V
by: Vergos, Polykarpos, et al.
Published: (2025)
by: Vergos, Polykarpos, et al.
Published: (2025)
SAT-based Exact Modulo Scheduling Mapping for Resource-Constrained CGRAs
by: Tirelli, Cristian, et al.
Published: (2024)
by: Tirelli, Cristian, et al.
Published: (2024)
Optimizing Layer-Fused Scheduling of Transformer Networks on Multi-accelerator Platforms
by: Colleman, Steven, et al.
Published: (2024)
by: Colleman, Steven, et al.
Published: (2024)
Retrieve, Schedule, Reflect: LLM Agents for Chip QoR Optimization
by: ouyang, Yikang, et al.
Published: (2026)
by: ouyang, Yikang, et al.
Published: (2026)
Multi-Dimensional Vector ISA Extension for Mobile In-Cache Computing
by: Khadem, Alireza, et al.
Published: (2025)
by: Khadem, Alireza, et al.
Published: (2025)
EONSim: An NPU Simulator for On-Chip Memory and Embedding Vector Operations
by: Choi, Sangun, et al.
Published: (2025)
by: Choi, Sangun, et al.
Published: (2025)
Empowering Vector Architectures for ML: The CAMP Architecture for Matrix Multiplication
by: Nojehdeh, Mohammadreza Esmali, et al.
Published: (2025)
by: Nojehdeh, Mohammadreza Esmali, et al.
Published: (2025)
PULSE: Parametric Hardware Units for Low-power Sparsity-Aware Convolution Engine
by: Aliyev, Ilkin, et al.
Published: (2024)
by: Aliyev, Ilkin, et al.
Published: (2024)
Virgo: Cluster-level Matrix Unit Integration in GPUs for Scalability and Energy Efficiency
by: Kim, Hansung, et al.
Published: (2024)
by: Kim, Hansung, et al.
Published: (2024)
Asynchronous Memory Access Unit: Exploiting Massive Parallelism for Far Memory Access
by: Wang, Luming, et al.
Published: (2024)
by: Wang, Luming, et al.
Published: (2024)
Accelerating Data Chunking in Deduplication Systems using Vector Instructions
by: Udayashankar, Sreeharsha, et al.
Published: (2025)
by: Udayashankar, Sreeharsha, et al.
Published: (2025)
Unlimited Vector Processing for Wireless Baseband Based on RISC-V Extension
by: Jiang, Limin, et al.
Published: (2025)
by: Jiang, Limin, et al.
Published: (2025)
Similar Items
-
NeCTAr: A Heterogeneous RISC-V SoC for Language Model Inference in Intel 16
by: Schmulbach, Viansa, et al.
Published: (2025) -
Efficient Implementation of RISC-V Vector Permutation Instructions
by: Titopoulos, Vasileios, et al.
Published: (2025) -
AraOS: Analyzing the Impact of Virtual Memory Management on Vector Unit Performance
by: Perotti, Matteo, et al.
Published: (2025) -
Instruction-Based Coordination of Heterogeneous Processing Units for Acceleration of DNN Inference
by: Petropoulos, Anastasios, et al.
Published: (2025) -
Spatz: Clustering Compact RISC-V-Based Vector Units to Maximize Computing Efficiency
by: Perotti, Matteo, et al.
Published: (2023)