Towards Cheaper Inference in Deep Networks with Lower Bit-Width Accumulators
Fuente:
arXiv
Salvato in:
| Autori principali: | Blumenfeld, Yaniv, Hubara, Itay, Soudry, Daniel |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LUT-DLA: Lookup Table as Efficient Extreme Low-Bit Deep Learning Accelerator
di: Li, Guoyu, et al.
Pubblicazione: (2025)
di: Li, Guoyu, et al.
Pubblicazione: (2025)
AMPLE: Event-Driven Accelerator for Mixed-Precision Inference of Graph Neural Networks
di: Gimenes, Pedro, et al.
Pubblicazione: (2025)
di: Gimenes, Pedro, et al.
Pubblicazione: (2025)
NeuralMatrix: Compute the Entire Neural Networks with Linear Matrix Operations for Efficient Inference
di: Sun, Ruiqi, et al.
Pubblicazione: (2023)
di: Sun, Ruiqi, et al.
Pubblicazione: (2023)
Estimation of Energy-dissipation Lower-bounds for Neuromorphic Learning-in-memory
di: Chen, Zihao, et al.
Pubblicazione: (2024)
di: Chen, Zihao, et al.
Pubblicazione: (2024)
LUTMUL: Exceed Conventional FPGA Roofline Limit by LUT-based Efficient Multiplication for Neural Network Inference
di: Xie, Yanyue, et al.
Pubblicazione: (2024)
di: Xie, Yanyue, et al.
Pubblicazione: (2024)
Deep Learning-Based Anomaly Detection in Spacecraft Telemetry on Edge Devices
di: Goetze, Christopher, et al.
Pubblicazione: (2026)
di: Goetze, Christopher, et al.
Pubblicazione: (2026)
RESQ: A Unified Framework for REliability- and Security Enhancement of Quantized Deep Neural Networks
di: Mohammadi, Ali Soltan, et al.
Pubblicazione: (2026)
di: Mohammadi, Ali Soltan, et al.
Pubblicazione: (2026)
HiFloat4 Format for Language Model Inference
di: Luo, Yuanyong, et al.
Pubblicazione: (2026)
di: Luo, Yuanyong, et al.
Pubblicazione: (2026)
Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference
di: Skliar, Andrii, et al.
Pubblicazione: (2024)
di: Skliar, Andrii, et al.
Pubblicazione: (2024)
Challenges and Research Directions for Large Language Model Inference Hardware
di: Ma, Xiaoyu, et al.
Pubblicazione: (2026)
di: Ma, Xiaoyu, et al.
Pubblicazione: (2026)
Continuous-Flow Data-Rate-Aware CNN Inference on FPGA
di: Habermann, Tobias, et al.
Pubblicazione: (2026)
di: Habermann, Tobias, et al.
Pubblicazione: (2026)
Runtime Tunable Tsetlin Machines for Edge Inference on eFPGAs
di: Rahman, Tousif, et al.
Pubblicazione: (2025)
di: Rahman, Tousif, et al.
Pubblicazione: (2025)
MixPE: Quantization and Hardware Co-design for Efficient LLM Inference
di: Zhang, Yu, et al.
Pubblicazione: (2024)
di: Zhang, Yu, et al.
Pubblicazione: (2024)
PolyLUT-Add: FPGA-based LUT Inference with Wide Inputs
di: Lou, Binglei, et al.
Pubblicazione: (2024)
di: Lou, Binglei, et al.
Pubblicazione: (2024)
Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns
di: Bambhaniya, Abhimanyu, et al.
Pubblicazione: (2026)
di: Bambhaniya, Abhimanyu, et al.
Pubblicazione: (2026)
HLSTransform: Energy-Efficient Llama 2 Inference on FPGAs Via High Level Synthesis
di: He, Andy, et al.
Pubblicazione: (2024)
di: He, Andy, et al.
Pubblicazione: (2024)
Hybrid JIT-CUDA Graph Optimization for Low-Latency Large Language Model Inference
di: Yadav, Divakar Kumar, et al.
Pubblicazione: (2026)
di: Yadav, Divakar Kumar, et al.
Pubblicazione: (2026)
ALADIN: Accuracy-Latency-Aware Design-space Inference Analysis for Embedded AI Accelerators
di: Baldi, T., et al.
Pubblicazione: (2026)
di: Baldi, T., et al.
Pubblicazione: (2026)
eIQ Neutron: Redefining Edge-AI Inference with Integrated NPU and Compiler Innovations
di: Bamberg, Lennart, et al.
Pubblicazione: (2025)
di: Bamberg, Lennart, et al.
Pubblicazione: (2025)
TRINE: A Token-Aware, Runtime-Adaptive FPGA Inference Engine for Multimodal AI
di: Oh, Hyunwoo, et al.
Pubblicazione: (2026)
di: Oh, Hyunwoo, et al.
Pubblicazione: (2026)
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format
di: Fang, Chao, et al.
Pubblicazione: (2024)
di: Fang, Chao, et al.
Pubblicazione: (2024)
AccLLM: Accelerating Long-Context LLM Inference Via Algorithm-Hardware Co-Design
di: Liang, Yanbiao, et al.
Pubblicazione: (2025)
di: Liang, Yanbiao, et al.
Pubblicazione: (2025)
Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference
di: Hwang, Ranggi, et al.
Pubblicazione: (2023)
di: Hwang, Ranggi, et al.
Pubblicazione: (2023)
Ascend HiFloat8 Format for Deep Learning
di: Luo, Yuanyong, et al.
Pubblicazione: (2024)
di: Luo, Yuanyong, et al.
Pubblicazione: (2024)
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
di: Ramachandran, Akshat, et al.
Pubblicazione: (2025)
di: Ramachandran, Akshat, et al.
Pubblicazione: (2025)
A Hybrid Edge Classifier: Combining TinyML-Optimised CNN with RRAM-CMOS ACAM for Energy-Efficient Inference
di: Woodward, Kieran, et al.
Pubblicazione: (2025)
di: Woodward, Kieran, et al.
Pubblicazione: (2025)
FPGA Divide-and-Conquer Placement using Deep Reinforcement Learning
di: Wang, Shang, et al.
Pubblicazione: (2024)
di: Wang, Shang, et al.
Pubblicazione: (2024)
Automatic Generation of Fast and Accurate Performance Models for Deep Neural Network Accelerators
di: Lübeck, Konstantin, et al.
Pubblicazione: (2024)
di: Lübeck, Konstantin, et al.
Pubblicazione: (2024)
Efficient Reprogramming of Memristive Crossbars for DNNs: Weight Sorting and Bit Stucking
di: Farias, Matheus, et al.
Pubblicazione: (2024)
di: Farias, Matheus, et al.
Pubblicazione: (2024)
SwiftRL: Towards Efficient Reinforcement Learning on Real Processing-In-Memory Systems
di: Gogineni, Kailash, et al.
Pubblicazione: (2024)
di: Gogineni, Kailash, et al.
Pubblicazione: (2024)
ElasticAI: Creating and Deploying Energy-Efficient Deep Learning Accelerator for Pervasive Computing
di: Qian, Chao, et al.
Pubblicazione: (2024)
di: Qian, Chao, et al.
Pubblicazione: (2024)
MG-Verilog: Multi-grained Dataset Towards Enhanced LLM-assisted Verilog Generation
di: Zhang, Yongan, et al.
Pubblicazione: (2024)
di: Zhang, Yongan, et al.
Pubblicazione: (2024)
A 10.60 $μ$W 150 GOPS Mixed-Bit-Width Sparse CNN Accelerator for Life-Threatening Ventricular Arrhythmia Detection
di: Qin, Yifan, et al.
Pubblicazione: (2024)
di: Qin, Yifan, et al.
Pubblicazione: (2024)
On-Sensor Convolutional Neural Networks with Early-Exits
di: Shalby, Hazem Hesham Yousef, et al.
Pubblicazione: (2025)
di: Shalby, Hazem Hesham Yousef, et al.
Pubblicazione: (2025)
eXmY: A Data Type and Technique for Arbitrary Bit Precision Quantization
di: Agrawal, Aditya, et al.
Pubblicazione: (2024)
di: Agrawal, Aditya, et al.
Pubblicazione: (2024)
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
di: Huang, Wei, et al.
Pubblicazione: (2023)
di: Huang, Wei, et al.
Pubblicazione: (2023)
GATMesh: Clock Mesh Timing Analysis using Graph Neural Networks
di: Khan, Muhammad Hadir, et al.
Pubblicazione: (2025)
di: Khan, Muhammad Hadir, et al.
Pubblicazione: (2025)
Revolutionizing TCAD Simulations with Universal Device Encoding and Graph Attention Networks
di: Fan, Guangxi, et al.
Pubblicazione: (2023)
di: Fan, Guangxi, et al.
Pubblicazione: (2023)
KirchhoffNet: A Scalable Ultra Fast Analog Neural Network
di: Gao, Zhengqi, et al.
Pubblicazione: (2023)
di: Gao, Zhengqi, et al.
Pubblicazione: (2023)
Automated and Holistic Co-design of Neural Networks and ASICs for Enabling In-Pixel Intelligence
di: Kharel, Shubha R., et al.
Pubblicazione: (2024)
di: Kharel, Shubha R., et al.
Pubblicazione: (2024)
Documenti analoghi
-
LUT-DLA: Lookup Table as Efficient Extreme Low-Bit Deep Learning Accelerator
di: Li, Guoyu, et al.
Pubblicazione: (2025) -
AMPLE: Event-Driven Accelerator for Mixed-Precision Inference of Graph Neural Networks
di: Gimenes, Pedro, et al.
Pubblicazione: (2025) -
NeuralMatrix: Compute the Entire Neural Networks with Linear Matrix Operations for Efficient Inference
di: Sun, Ruiqi, et al.
Pubblicazione: (2023) -
Estimation of Energy-dissipation Lower-bounds for Neuromorphic Learning-in-memory
di: Chen, Zihao, et al.
Pubblicazione: (2024) -
LUTMUL: Exceed Conventional FPGA Roofline Limit by LUT-based Efficient Multiplication for Neural Network Inference
di: Xie, Yanyue, et al.
Pubblicazione: (2024)