Beehive: A Flexible Network Stack for Direct-Attached Accelerators
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lim, Katie, Giordano, Matthew, Stavrinos, Theano, Zhang, Irene, Nelson, Jacob, Kasikci, Baris, Anderson, Tom |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FETTA: Flexible and Efficient Hardware Accelerator for Tensorized Neural Network Training
von: Lu, Jinming, et al.
Veröffentlicht: (2025)
von: Lu, Jinming, et al.
Veröffentlicht: (2025)
Arcus: SLO Management for Accelerators in the Cloud with Traffic Shaping
von: Zhao, Jiechen, et al.
Veröffentlicht: (2024)
von: Zhao, Jiechen, et al.
Veröffentlicht: (2024)
Architecture, Simulation and Software Stack to Support Post-CMOS Accelerators: The ARCHYTAS Project
von: Agosta, Giovanni, et al.
Veröffentlicht: (2025)
von: Agosta, Giovanni, et al.
Veröffentlicht: (2025)
A Full-Stack Performance Evaluation Infrastructure for 3D-DRAM-based LLM Accelerators
von: Li, Cong, et al.
Veröffentlicht: (2026)
von: Li, Cong, et al.
Veröffentlicht: (2026)
GPU-Accelerated Simulated Oscillator Ising/Potts Machine Solving Combinatorial Optimization Problems
von: Gonul, Yilmaz Ege, et al.
Veröffentlicht: (2025)
von: Gonul, Yilmaz Ege, et al.
Veröffentlicht: (2025)
A Multi-Stage Potts Machine based on Coupled CMOS Ring Oscillators
von: Gonul, Yilmaz Ege, et al.
Veröffentlicht: (2025)
von: Gonul, Yilmaz Ege, et al.
Veröffentlicht: (2025)
FILCO: Flexible Composing Architecture with Real-Time Reconfigurability for DNN Acceleration
von: Chen, Xingzhen, et al.
Veröffentlicht: (2026)
von: Chen, Xingzhen, et al.
Veröffentlicht: (2026)
A Flexible Template for Edge Generative AI with High-Accuracy Accelerated Softmax & GELU
von: Belano, Andrea, et al.
Veröffentlicht: (2024)
von: Belano, Andrea, et al.
Veröffentlicht: (2024)
Bancroft: Genomics Acceleration Beyond On-Device Memory
von: Lim, Se-Min, et al.
Veröffentlicht: (2025)
von: Lim, Se-Min, et al.
Veröffentlicht: (2025)
EFFACT: A Highly Efficient Full-Stack FHE Acceleration Platform
von: Huang, Yi, et al.
Veröffentlicht: (2025)
von: Huang, Yi, et al.
Veröffentlicht: (2025)
3D Stack In-Sensor-Computing (3DS-ISC): Accelerating Time-Surface Construction for Neuromorphic Event Cameras
von: Shang, Hongyang, et al.
Veröffentlicht: (2025)
von: Shang, Hongyang, et al.
Veröffentlicht: (2025)
A Switch-Centric In-Network Architecture for Accelerating LLM Inference in Shared-Memory Network
von: Jiang, Aojie, et al.
Veröffentlicht: (2026)
von: Jiang, Aojie, et al.
Veröffentlicht: (2026)
Palermo: Improving the Performance of Oblivious Memory using Protocol-Hardware Co-Design
von: Ye, Haojie, et al.
Veröffentlicht: (2024)
von: Ye, Haojie, et al.
Veröffentlicht: (2024)
A Flexible Precision Scaling Deep Neural Network Accelerator with Efficient Weight Combination
von: Zhao, Liang, et al.
Veröffentlicht: (2025)
von: Zhao, Liang, et al.
Veröffentlicht: (2025)
All-rounder: A Flexible AI Accelerator with Diverse Data Format Support and Morphable Structure for Multi-DNN Processing
von: Noh, Seock-Hwan, et al.
Veröffentlicht: (2023)
von: Noh, Seock-Hwan, et al.
Veröffentlicht: (2023)
31.1 A 14.08-to-135.69Token/s ReRAM-on-Logic Stacked Outlier-Free Large-Language-Model Accelerator with Block-Clustered Weight-Compression and Adaptive Parallel-Speculative-Decoding
von: Dong, Pingcheng, et al.
Veröffentlicht: (2026)
von: Dong, Pingcheng, et al.
Veröffentlicht: (2026)
MXFormer: A Microscaling Floating-Point Charge-Trap Transistor Compute-in-Memory Transformer Accelerator
von: Karfakis, George, et al.
Veröffentlicht: (2026)
von: Karfakis, George, et al.
Veröffentlicht: (2026)
When Pipelined In-Memory Accelerators Meet Spiking Direct Feedback Alignment: A Co-Design for Neuromorphic Edge Computing
von: Ren, Haoxiong, et al.
Veröffentlicht: (2025)
von: Ren, Haoxiong, et al.
Veröffentlicht: (2025)
FastFlow in FPGA Stacks of Data Centers
von: Paul, Rourab, et al.
Veröffentlicht: (2024)
von: Paul, Rourab, et al.
Veröffentlicht: (2024)
TensorPool: A 3D-Stacked 8.4TFLOPS/4.3W Many-Core Domain-Specific Processor for AI-Native Radio Access Networks
von: Bertuletti, Marco, et al.
Veröffentlicht: (2026)
von: Bertuletti, Marco, et al.
Veröffentlicht: (2026)
Prosperity: Accelerating Spiking Neural Networks via Product Sparsity
von: Wei, Chiyue, et al.
Veröffentlicht: (2025)
von: Wei, Chiyue, et al.
Veröffentlicht: (2025)
FireFly-P: FPGA-Accelerated Spiking Neural Network Plasticity for Robust Adaptive Control
von: Li, Tenglong, et al.
Veröffentlicht: (2026)
von: Li, Tenglong, et al.
Veröffentlicht: (2026)
Modeling and Optimizing Performance Bottlenecks for Neuromorphic Accelerators
von: Yik, Jason, et al.
Veröffentlicht: (2025)
von: Yik, Jason, et al.
Veröffentlicht: (2025)
FireFly-S: Exploiting Dual-Side Sparsity for Spiking Neural Networks Acceleration with Reconfigurable Spatial Architecture
von: Li, Tenglong, et al.
Veröffentlicht: (2024)
von: Li, Tenglong, et al.
Veröffentlicht: (2024)
Hardware Acceleration of Kolmogorov-Arnold Network (KAN) for Lightweight Edge Inference
von: Huang, Wei-Hsing, et al.
Veröffentlicht: (2024)
von: Huang, Wei-Hsing, et al.
Veröffentlicht: (2024)
Hardware Acceleration of Kolmogorov-Arnold Network (KAN) in Large-Scale Systems
von: Huang, Wei-Hsing, et al.
Veröffentlicht: (2025)
von: Huang, Wei-Hsing, et al.
Veröffentlicht: (2025)
CHIMERA: A Flexible and Scalable 3.1 TOPS/W AI-MCU with Transformer Accelerator and 563 Gb/s Shared-L2 Memory Subsystem with QoS Guarantees
von: Leone, Lorenzo, et al.
Veröffentlicht: (2026)
von: Leone, Lorenzo, et al.
Veröffentlicht: (2026)
GDR-HGNN: A Heterogeneous Graph Neural Networks Accelerator Frontend with Graph Decoupling and Recoupling
von: Xue, Runzhen, et al.
Veröffentlicht: (2024)
von: Xue, Runzhen, et al.
Veröffentlicht: (2024)
Hardware-Aware Neural Network Compilation with Learned Optimization: A RISC-V Accelerator Approach
von: Ganti, Ravindra, et al.
Veröffentlicht: (2025)
von: Ganti, Ravindra, et al.
Veröffentlicht: (2025)
FastCaps: A Design Methodology for Accelerating Capsule Network on Field Programmable Gate Arrays
von: Rahoof, Abdul, et al.
Veröffentlicht: (2025)
von: Rahoof, Abdul, et al.
Veröffentlicht: (2025)
Optical Computing for Deep Neural Network Acceleration: Foundations, Recent Developments, and Emerging Directions
von: Pasricha, Sudeep
Veröffentlicht: (2024)
von: Pasricha, Sudeep
Veröffentlicht: (2024)
APINT: A Full-Stack Framework for Acceleration of Privacy-Preserving Inference of Transformers based on Garbled Circuits
von: Cho, Hyunjun, et al.
Veröffentlicht: (2025)
von: Cho, Hyunjun, et al.
Veröffentlicht: (2025)
Enabling Efficient Hardware Acceleration of Hybrid Vision Transformer (ViT) Networks at the Edge
von: Dumoulin, Joren, et al.
Veröffentlicht: (2025)
von: Dumoulin, Joren, et al.
Veröffentlicht: (2025)
CIMPool: Scalable Neural Network Acceleration for Compute-In-Memory using Weight Pools
von: Li, Shurui, et al.
Veröffentlicht: (2025)
von: Li, Shurui, et al.
Veröffentlicht: (2025)
Understanding Simulated Architecture via gem5 Call-Stack Profiling
von: Söderström, Johan, et al.
Veröffentlicht: (2026)
von: Söderström, Johan, et al.
Veröffentlicht: (2026)
Rethinking Compute Substrates for 3D-Stacked Near-Memory LLM Decoding: Microarchitecture-Scheduling Co-Design
von: Ai, Chenyang, et al.
Veröffentlicht: (2026)
von: Ai, Chenyang, et al.
Veröffentlicht: (2026)
SIRA: Scaled-Integer Range Analysis for Optimizing FPGA Dataflow Neural Network Accelerators
von: Umuroglu, Yaman, et al.
Veröffentlicht: (2025)
von: Umuroglu, Yaman, et al.
Veröffentlicht: (2025)
Fast-OverlaPIM: A Fast Overlap-driven Mapping Framework for Processing In-Memory Neural Network Acceleration
von: Wang, Xuan, et al.
Veröffentlicht: (2024)
von: Wang, Xuan, et al.
Veröffentlicht: (2024)
ApproxPilot: A GNN-based Accelerator Approximation Framework
von: Zhang, Qing, et al.
Veröffentlicht: (2024)
von: Zhang, Qing, et al.
Veröffentlicht: (2024)
Computing with Printed and Flexible Electronics
von: Tahoori, Mehdi B., et al.
Veröffentlicht: (2025)
von: Tahoori, Mehdi B., et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FETTA: Flexible and Efficient Hardware Accelerator for Tensorized Neural Network Training
von: Lu, Jinming, et al.
Veröffentlicht: (2025) -
Arcus: SLO Management for Accelerators in the Cloud with Traffic Shaping
von: Zhao, Jiechen, et al.
Veröffentlicht: (2024) -
Architecture, Simulation and Software Stack to Support Post-CMOS Accelerators: The ARCHYTAS Project
von: Agosta, Giovanni, et al.
Veröffentlicht: (2025) -
A Full-Stack Performance Evaluation Infrastructure for 3D-DRAM-based LLM Accelerators
von: Li, Cong, et al.
Veröffentlicht: (2026) -
GPU-Accelerated Simulated Oscillator Ising/Potts Machine Solving Combinatorial Optimization Problems
von: Gonul, Yilmaz Ege, et al.
Veröffentlicht: (2025)