HADES: Hardware Accelerated Decoding for Efficient Speculation in Large Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yang, Ze, Jin, Yihong, Xu, Xinhe |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Hardware Phi-1.5B: A Large Language Model Encodes Hardware Domain Specific Knowledge
par: Fu, Weimin, et autres
Publié: (2024)
par: Fu, Weimin, et autres
Publié: (2024)
From English to ASIC: Hardware Implementation with Large Language Model
par: Goh, Emil, et autres
Publié: (2024)
par: Goh, Emil, et autres
Publié: (2024)
LLM Inference Acceleration via Efficient Operation Fusion
par: Salmani, Mahsa, et autres
Publié: (2025)
par: Salmani, Mahsa, et autres
Publié: (2025)
Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding
par: Fan, Wang, et autres
Publié: (2026)
par: Fan, Wang, et autres
Publié: (2026)
DocEDA: Automated Extraction and Design of Analog Circuits from Documents with Large Language Model
par: Chen, Hong Cai, et autres
Publié: (2024)
par: Chen, Hong Cai, et autres
Publié: (2024)
Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inference
par: Chen, Hongzheng, et autres
Publié: (2023)
par: Chen, Hongzheng, et autres
Publié: (2023)
EDA Corpus: A Large Language Model Dataset for Enhanced Interaction with OpenROAD
par: Wu, Bing-Yue, et autres
Publié: (2024)
par: Wu, Bing-Yue, et autres
Publié: (2024)
LLM-FSM: Scaling Large Language Models for Finite-State Reasoning in RTL Code Generation
par: Wu, Yuheng, et autres
Publié: (2026)
par: Wu, Yuheng, et autres
Publié: (2026)
Using the Abstract Computer Architecture Description Language to Model AI Hardware Accelerators
par: Müller, Mika Markus, et autres
Publié: (2024)
par: Müller, Mika Markus, et autres
Publié: (2024)
A Configurable and Efficient Memory Hierarchy for Neural Network Hardware Accelerator
par: Bause, Oliver, et autres
Publié: (2024)
par: Bause, Oliver, et autres
Publié: (2024)
Chameleon: a Heterogeneous and Disaggregated Accelerator System for Retrieval-Augmented Language Models
par: Jiang, Wenqi, et autres
Publié: (2023)
par: Jiang, Wenqi, et autres
Publié: (2023)
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
par: Du, Dayou, et autres
Publié: (2025)
par: Du, Dayou, et autres
Publié: (2025)
LLM4SecHW: Leveraging Domain Specific Large Language Model for Hardware Debugging
par: Fu, Weimin, et autres
Publié: (2024)
par: Fu, Weimin, et autres
Publié: (2024)
EdgeCIM: A Hardware-Software Co-Design for CIM-Based Acceleration of Small Language Models
par: Bazzi, Jinane, et autres
Publié: (2026)
par: Bazzi, Jinane, et autres
Publié: (2026)
Towards Efficient IMC Accelerator Design Through Joint Hardware-Workload Co-optimization
par: Krestinskaya, Olga, et autres
Publié: (2024)
par: Krestinskaya, Olga, et autres
Publié: (2024)
ApproXAI: Energy-Efficient Hardware Acceleration of Explainable AI using Approximate Computing
par: Siddique, Ayesha, et autres
Publié: (2025)
par: Siddique, Ayesha, et autres
Publié: (2025)
InCoder-32B-Thinking: Industrial Code World Model for Thinking
par: Yang, Jian, et autres
Publié: (2026)
par: Yang, Jian, et autres
Publié: (2026)
Speculative Decoding for Verilog: Speed and Quality, All in One
par: Xu, Changran, et autres
Publié: (2025)
par: Xu, Changran, et autres
Publié: (2025)
AHASD: Asynchronous Heterogeneous Architecture for LLM Adaptive Drafting Speculative Decoding on Mobile Devices
par: Zirui, Ma, et autres
Publié: (2026)
par: Zirui, Ma, et autres
Publié: (2026)
Customizing a Large Language Model for VHDL Design of High-Performance Microprocessors
par: Dupuis, Nicolas, et autres
Publié: (2025)
par: Dupuis, Nicolas, et autres
Publié: (2025)
QiMeng-CodeV-SVA: Training Specialized LLMs for Hardware Assertion Generation via RTL-Grounded Bidirectional Data Synthesis
par: Wu, Yutong, et autres
Publié: (2026)
par: Wu, Yutong, et autres
Publié: (2026)
Revisiting VerilogEval: A Year of Improvements in Large-Language Models for Hardware Code Generation
par: Pinckney, Nathaniel, et autres
Publié: (2024)
par: Pinckney, Nathaniel, et autres
Publié: (2024)
Hardware Acceleration of LLMs: A comprehensive survey and comparison
par: Koilia, Nikoletta, et autres
Publié: (2024)
par: Koilia, Nikoletta, et autres
Publié: (2024)
White-Box Reasoning: Synergizing LLM Strategy and gm/Id Data for Automated Analog Circuit Design
par: Chen, Jianqiu, et autres
Publié: (2025)
par: Chen, Jianqiu, et autres
Publié: (2025)
Evaluating LLMs for Hardware Design and Test
par: Blocklove, Jason, et autres
Publié: (2024)
par: Blocklove, Jason, et autres
Publié: (2024)
HaLoRA: Hardware-aware Low-Rank Adaptation for Large Language Models Based on Hybrid Compute-in-Memory Architecture
par: Wu, Taiqiang, et autres
Publié: (2025)
par: Wu, Taiqiang, et autres
Publié: (2025)
Comprehensive Design Space Exploration for Tensorized Neural Network Hardware Accelerators
par: Zhang, Jinsong, et autres
Publié: (2025)
par: Zhang, Jinsong, et autres
Publié: (2025)
D2S-FLOW: Automated Parameter Extraction from Datasheets for SPICE Model Generation Using Large Language Models
par: Chen, Hong Cai, et autres
Publié: (2025)
par: Chen, Hong Cai, et autres
Publié: (2025)
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference
par: Chen, Chun-Ting, et autres
Publié: (2025)
par: Chen, Chun-Ting, et autres
Publié: (2025)
FVEval: Understanding Language Model Capabilities in Formal Verification of Digital Hardware
par: Kang, Minwoo, et autres
Publié: (2024)
par: Kang, Minwoo, et autres
Publié: (2024)
GRAU: Generic Reconfigurable Activation Unit Design for Neural Network Hardware Accelerators
par: Liu, Yuhao, et autres
Publié: (2026)
par: Liu, Yuhao, et autres
Publié: (2026)
LUT-LLM: Efficient Large Language Model Inference with Memory-based Computations on FPGAs
par: He, Zifan, et autres
Publié: (2025)
par: He, Zifan, et autres
Publié: (2025)
SpikeX: Exploring Accelerator Architecture and Network-Hardware Co-Optimization for Sparse Spiking Neural Networks
par: Xu, Boxun, et autres
Publié: (2025)
par: Xu, Boxun, et autres
Publié: (2025)
SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators
par: Li, Jonathan, et autres
Publié: (2025)
par: Li, Jonathan, et autres
Publié: (2025)
Accelerating Post-Quantum Cryptography via LLM-Driven Hardware-Software Co-Design
par: Liao, Yuchao, et autres
Publié: (2026)
par: Liao, Yuchao, et autres
Publié: (2026)
Digital ASIC Design with Ongoing LLMs: Strategies and Prospects
par: Xiang, Maoyang, et autres
Publié: (2024)
par: Xiang, Maoyang, et autres
Publié: (2024)
Chain-of-Descriptions: Improving Code LLMs for VHDL Code Generation and Summarization
par: Vijayaraghavan, Prashanth, et autres
Publié: (2025)
par: Vijayaraghavan, Prashanth, et autres
Publié: (2025)
ProtocolLLM: RTL Benchmark for SystemVerilog Generation of Communication Protocols
par: Sheth, Arnav, et autres
Publié: (2025)
par: Sheth, Arnav, et autres
Publié: (2025)
FloorPlan-DeepSeek (FPDS): A multimodal approach to floorplan generation using vector-based next room prediction
par: Yin, Jun, et autres
Publié: (2025)
par: Yin, Jun, et autres
Publié: (2025)
A Fully Hardware Implemented Accelerator Design in ReRAM Analog Computing without ADCs
par: Dang, Peng, et autres
Publié: (2024)
par: Dang, Peng, et autres
Publié: (2024)
Documents similaires
-
Hardware Phi-1.5B: A Large Language Model Encodes Hardware Domain Specific Knowledge
par: Fu, Weimin, et autres
Publié: (2024) -
From English to ASIC: Hardware Implementation with Large Language Model
par: Goh, Emil, et autres
Publié: (2024) -
LLM Inference Acceleration via Efficient Operation Fusion
par: Salmani, Mahsa, et autres
Publié: (2025) -
Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding
par: Fan, Wang, et autres
Publié: (2026) -
DocEDA: Automated Extraction and Design of Analog Circuits from Documents with Large Language Model
par: Chen, Hong Cai, et autres
Publié: (2024)