Beehive: A Flexible Network Stack for Direct-Attached Accelerators
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Lim, Katie, Giordano, Matthew, Stavrinos, Theano, Zhang, Irene, Nelson, Jacob, Kasikci, Baris, Anderson, Tom |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
FETTA: Flexible and Efficient Hardware Accelerator for Tensorized Neural Network Training
par: Lu, Jinming, et autres
Publié: (2025)
par: Lu, Jinming, et autres
Publié: (2025)
Arcus: SLO Management for Accelerators in the Cloud with Traffic Shaping
par: Zhao, Jiechen, et autres
Publié: (2024)
par: Zhao, Jiechen, et autres
Publié: (2024)
Architecture, Simulation and Software Stack to Support Post-CMOS Accelerators: The ARCHYTAS Project
par: Agosta, Giovanni, et autres
Publié: (2025)
par: Agosta, Giovanni, et autres
Publié: (2025)
A Full-Stack Performance Evaluation Infrastructure for 3D-DRAM-based LLM Accelerators
par: Li, Cong, et autres
Publié: (2026)
par: Li, Cong, et autres
Publié: (2026)
GPU-Accelerated Simulated Oscillator Ising/Potts Machine Solving Combinatorial Optimization Problems
par: Gonul, Yilmaz Ege, et autres
Publié: (2025)
par: Gonul, Yilmaz Ege, et autres
Publié: (2025)
A Multi-Stage Potts Machine based on Coupled CMOS Ring Oscillators
par: Gonul, Yilmaz Ege, et autres
Publié: (2025)
par: Gonul, Yilmaz Ege, et autres
Publié: (2025)
FILCO: Flexible Composing Architecture with Real-Time Reconfigurability for DNN Acceleration
par: Chen, Xingzhen, et autres
Publié: (2026)
par: Chen, Xingzhen, et autres
Publié: (2026)
A Flexible Template for Edge Generative AI with High-Accuracy Accelerated Softmax & GELU
par: Belano, Andrea, et autres
Publié: (2024)
par: Belano, Andrea, et autres
Publié: (2024)
Bancroft: Genomics Acceleration Beyond On-Device Memory
par: Lim, Se-Min, et autres
Publié: (2025)
par: Lim, Se-Min, et autres
Publié: (2025)
EFFACT: A Highly Efficient Full-Stack FHE Acceleration Platform
par: Huang, Yi, et autres
Publié: (2025)
par: Huang, Yi, et autres
Publié: (2025)
3D Stack In-Sensor-Computing (3DS-ISC): Accelerating Time-Surface Construction for Neuromorphic Event Cameras
par: Shang, Hongyang, et autres
Publié: (2025)
par: Shang, Hongyang, et autres
Publié: (2025)
A Switch-Centric In-Network Architecture for Accelerating LLM Inference in Shared-Memory Network
par: Jiang, Aojie, et autres
Publié: (2026)
par: Jiang, Aojie, et autres
Publié: (2026)
Palermo: Improving the Performance of Oblivious Memory using Protocol-Hardware Co-Design
par: Ye, Haojie, et autres
Publié: (2024)
par: Ye, Haojie, et autres
Publié: (2024)
A Flexible Precision Scaling Deep Neural Network Accelerator with Efficient Weight Combination
par: Zhao, Liang, et autres
Publié: (2025)
par: Zhao, Liang, et autres
Publié: (2025)
All-rounder: A Flexible AI Accelerator with Diverse Data Format Support and Morphable Structure for Multi-DNN Processing
par: Noh, Seock-Hwan, et autres
Publié: (2023)
par: Noh, Seock-Hwan, et autres
Publié: (2023)
31.1 A 14.08-to-135.69Token/s ReRAM-on-Logic Stacked Outlier-Free Large-Language-Model Accelerator with Block-Clustered Weight-Compression and Adaptive Parallel-Speculative-Decoding
par: Dong, Pingcheng, et autres
Publié: (2026)
par: Dong, Pingcheng, et autres
Publié: (2026)
MXFormer: A Microscaling Floating-Point Charge-Trap Transistor Compute-in-Memory Transformer Accelerator
par: Karfakis, George, et autres
Publié: (2026)
par: Karfakis, George, et autres
Publié: (2026)
When Pipelined In-Memory Accelerators Meet Spiking Direct Feedback Alignment: A Co-Design for Neuromorphic Edge Computing
par: Ren, Haoxiong, et autres
Publié: (2025)
par: Ren, Haoxiong, et autres
Publié: (2025)
FastFlow in FPGA Stacks of Data Centers
par: Paul, Rourab, et autres
Publié: (2024)
par: Paul, Rourab, et autres
Publié: (2024)
TensorPool: A 3D-Stacked 8.4TFLOPS/4.3W Many-Core Domain-Specific Processor for AI-Native Radio Access Networks
par: Bertuletti, Marco, et autres
Publié: (2026)
par: Bertuletti, Marco, et autres
Publié: (2026)
Prosperity: Accelerating Spiking Neural Networks via Product Sparsity
par: Wei, Chiyue, et autres
Publié: (2025)
par: Wei, Chiyue, et autres
Publié: (2025)
FireFly-P: FPGA-Accelerated Spiking Neural Network Plasticity for Robust Adaptive Control
par: Li, Tenglong, et autres
Publié: (2026)
par: Li, Tenglong, et autres
Publié: (2026)
Modeling and Optimizing Performance Bottlenecks for Neuromorphic Accelerators
par: Yik, Jason, et autres
Publié: (2025)
par: Yik, Jason, et autres
Publié: (2025)
FireFly-S: Exploiting Dual-Side Sparsity for Spiking Neural Networks Acceleration with Reconfigurable Spatial Architecture
par: Li, Tenglong, et autres
Publié: (2024)
par: Li, Tenglong, et autres
Publié: (2024)
Hardware Acceleration of Kolmogorov-Arnold Network (KAN) for Lightweight Edge Inference
par: Huang, Wei-Hsing, et autres
Publié: (2024)
par: Huang, Wei-Hsing, et autres
Publié: (2024)
Hardware Acceleration of Kolmogorov-Arnold Network (KAN) in Large-Scale Systems
par: Huang, Wei-Hsing, et autres
Publié: (2025)
par: Huang, Wei-Hsing, et autres
Publié: (2025)
CHIMERA: A Flexible and Scalable 3.1 TOPS/W AI-MCU with Transformer Accelerator and 563 Gb/s Shared-L2 Memory Subsystem with QoS Guarantees
par: Leone, Lorenzo, et autres
Publié: (2026)
par: Leone, Lorenzo, et autres
Publié: (2026)
GDR-HGNN: A Heterogeneous Graph Neural Networks Accelerator Frontend with Graph Decoupling and Recoupling
par: Xue, Runzhen, et autres
Publié: (2024)
par: Xue, Runzhen, et autres
Publié: (2024)
Hardware-Aware Neural Network Compilation with Learned Optimization: A RISC-V Accelerator Approach
par: Ganti, Ravindra, et autres
Publié: (2025)
par: Ganti, Ravindra, et autres
Publié: (2025)
FastCaps: A Design Methodology for Accelerating Capsule Network on Field Programmable Gate Arrays
par: Rahoof, Abdul, et autres
Publié: (2025)
par: Rahoof, Abdul, et autres
Publié: (2025)
Optical Computing for Deep Neural Network Acceleration: Foundations, Recent Developments, and Emerging Directions
par: Pasricha, Sudeep
Publié: (2024)
par: Pasricha, Sudeep
Publié: (2024)
APINT: A Full-Stack Framework for Acceleration of Privacy-Preserving Inference of Transformers based on Garbled Circuits
par: Cho, Hyunjun, et autres
Publié: (2025)
par: Cho, Hyunjun, et autres
Publié: (2025)
Enabling Efficient Hardware Acceleration of Hybrid Vision Transformer (ViT) Networks at the Edge
par: Dumoulin, Joren, et autres
Publié: (2025)
par: Dumoulin, Joren, et autres
Publié: (2025)
CIMPool: Scalable Neural Network Acceleration for Compute-In-Memory using Weight Pools
par: Li, Shurui, et autres
Publié: (2025)
par: Li, Shurui, et autres
Publié: (2025)
Understanding Simulated Architecture via gem5 Call-Stack Profiling
par: Söderström, Johan, et autres
Publié: (2026)
par: Söderström, Johan, et autres
Publié: (2026)
Rethinking Compute Substrates for 3D-Stacked Near-Memory LLM Decoding: Microarchitecture-Scheduling Co-Design
par: Ai, Chenyang, et autres
Publié: (2026)
par: Ai, Chenyang, et autres
Publié: (2026)
SIRA: Scaled-Integer Range Analysis for Optimizing FPGA Dataflow Neural Network Accelerators
par: Umuroglu, Yaman, et autres
Publié: (2025)
par: Umuroglu, Yaman, et autres
Publié: (2025)
Fast-OverlaPIM: A Fast Overlap-driven Mapping Framework for Processing In-Memory Neural Network Acceleration
par: Wang, Xuan, et autres
Publié: (2024)
par: Wang, Xuan, et autres
Publié: (2024)
ApproxPilot: A GNN-based Accelerator Approximation Framework
par: Zhang, Qing, et autres
Publié: (2024)
par: Zhang, Qing, et autres
Publié: (2024)
Computing with Printed and Flexible Electronics
par: Tahoori, Mehdi B., et autres
Publié: (2025)
par: Tahoori, Mehdi B., et autres
Publié: (2025)
Documents similaires
-
FETTA: Flexible and Efficient Hardware Accelerator for Tensorized Neural Network Training
par: Lu, Jinming, et autres
Publié: (2025) -
Arcus: SLO Management for Accelerators in the Cloud with Traffic Shaping
par: Zhao, Jiechen, et autres
Publié: (2024) -
Architecture, Simulation and Software Stack to Support Post-CMOS Accelerators: The ARCHYTAS Project
par: Agosta, Giovanni, et autres
Publié: (2025) -
A Full-Stack Performance Evaluation Infrastructure for 3D-DRAM-based LLM Accelerators
par: Li, Cong, et autres
Publié: (2026) -
GPU-Accelerated Simulated Oscillator Ising/Potts Machine Solving Combinatorial Optimization Problems
par: Gonul, Yilmaz Ege, et autres
Publié: (2025)