From Quarter to All: Accelerating Speculative LLM Decoding via Floating-Point Exponent Remapping and Parameter Sharing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Yushu, Qin, Yubin, Wang, Yang, Yang, Xiaolong, Han, Huiming, Wei, Shaojun, Hu, Yang, Yin, Shouyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
SOFA: A Compute-Memory Optimized Sparsity Accelerator via Cross-Stage Coordinated Tiling
von: Wang, Huizheng, et al.
Veröffentlicht: (2024)
von: Wang, Huizheng, et al.
Veröffentlicht: (2024)
BBAL: A Bidirectional Block Floating Point-Based Quantisation Accelerator for Large Language Models
von: Han, Xiaomeng, et al.
Veröffentlicht: (2025)
von: Han, Xiaomeng, et al.
Veröffentlicht: (2025)
SpecMamba: Accelerating Mamba Inference on FPGA with Speculative Decoding
von: Zhong, Linfeng, et al.
Veröffentlicht: (2025)
von: Zhong, Linfeng, et al.
Veröffentlicht: (2025)
Designing Spatial Architectures for Sparse Attention: STAR Accelerator via Cross-Stage Tiling
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
HADES: Hardware Accelerated Decoding for Efficient Speculation in Large Language Models
von: Yang, Ze, et al.
Veröffentlicht: (2024)
von: Yang, Ze, et al.
Veröffentlicht: (2024)
H-FA: A Hybrid Floating-Point and Logarithmic Approach to Hardware Accelerated FlashAttention
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2025)
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2025)
MXFormer: A Microscaling Floating-Point Charge-Trap Transistor Compute-in-Memory Transformer Accelerator
von: Karfakis, George, et al.
Veröffentlicht: (2026)
von: Karfakis, George, et al.
Veröffentlicht: (2026)
TimeFloats: Train-in-Memory with Time-Domain Floating-Point Scalar Products
von: Hashem, Maeesha Binte, et al.
Veröffentlicht: (2024)
von: Hashem, Maeesha Binte, et al.
Veröffentlicht: (2024)
Procrastination Is All You Need: Exponent Indexed Accumulators for Floating Point, Posits and Logarithmic Numbers
von: Liguori, Vincenzo
Veröffentlicht: (2024)
von: Liguori, Vincenzo
Veröffentlicht: (2024)
PADE: A Predictor-Free Sparse Attention Accelerator via Unified Execution and Stage Fusion
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
The AetherFloat Family: Block-Scale-Free Quad-Radix Floating-Point Architectures for AI Accelerators
von: Morisaki, Keita
Veröffentlicht: (2026)
von: Morisaki, Keita
Veröffentlicht: (2026)
A Hybrid-Domain Floating-Point Compute-in-Memory Architecture for Efficient Acceleration of High-Precision Deep Neural Networks
von: Yi, Zhiqiang, et al.
Veröffentlicht: (2025)
von: Yi, Zhiqiang, et al.
Veröffentlicht: (2025)
Titanus: Enabling KV Cache Pruning and Quantization On-the-Fly for LLM Acceleration
von: Chen, Peilin, et al.
Veröffentlicht: (2025)
von: Chen, Peilin, et al.
Veröffentlicht: (2025)
Fast Generation of Custom Floating-Point Spatial Filters on FPGAs
von: Campos, Nelson, et al.
Veröffentlicht: (2024)
von: Campos, Nelson, et al.
Veröffentlicht: (2024)
Online Alignment and Addition in Multi-Term Floating-Point Adders
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2024)
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2024)
Floating Point HUB Adder for RISC-V Sargantana Processor
von: Bandera, Gerardo, et al.
Veröffentlicht: (2024)
von: Bandera, Gerardo, et al.
Veröffentlicht: (2024)
Efficient Orchestrated AI Workflows Execution on Scale-out Spatial Architecture
von: Deng, Jinyi, et al.
Veröffentlicht: (2024)
von: Deng, Jinyi, et al.
Veröffentlicht: (2024)
ISAAC: Intelligent, Scalable, Agile, and Accelerated CPU Verification via LLM-aided FPGA Parallelism
von: Sun, Jialin, et al.
Veröffentlicht: (2025)
von: Sun, Jialin, et al.
Veröffentlicht: (2025)
Speculative Decoding for Verilog: Speed and Quality, All in One
von: Xu, Changran, et al.
Veröffentlicht: (2025)
von: Xu, Changran, et al.
Veröffentlicht: (2025)
Accelerator-assisted Floating-point ASIP for Communication and Positioning in Massive MIMO Systems
von: Attari, Mohammad, et al.
Veröffentlicht: (2025)
von: Attari, Mohammad, et al.
Veröffentlicht: (2025)
Floating-Point Multiply-Add with Approximate Normalization for Low-Cost Matrix Engines
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2024)
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2024)
Converting Binary Floating-Point Numbers to Shortest Decimal Strings: An Experimental Review
von: Gareau, Jaël Champagne, et al.
Veröffentlicht: (2026)
von: Gareau, Jaël Champagne, et al.
Veröffentlicht: (2026)
31.1 A 14.08-to-135.69Token/s ReRAM-on-Logic Stacked Outlier-Free Large-Language-Model Accelerator with Block-Clustered Weight-Compression and Adaptive Parallel-Speculative-Decoding
von: Dong, Pingcheng, et al.
Veröffentlicht: (2026)
von: Dong, Pingcheng, et al.
Veröffentlicht: (2026)
RTL Verification for Secure Speculation Using Contract Shadow Logic
von: Tan, Qinhan, et al.
Veröffentlicht: (2024)
von: Tan, Qinhan, et al.
Veröffentlicht: (2024)
A Switch-Centric In-Network Architecture for Accelerating LLM Inference in Shared-Memory Network
von: Jiang, Aojie, et al.
Veröffentlicht: (2026)
von: Jiang, Aojie, et al.
Veröffentlicht: (2026)
A Stochastic Rounding-Enabled Low-Precision Floating-Point MAC for DNN Training
von: Ali, Sami Ben, et al.
Veröffentlicht: (2024)
von: Ali, Sami Ben, et al.
Veröffentlicht: (2024)
TEMP: A Memory Efficient Physical-aware Tensor Partition-Mapping Framework on Wafer-scale Chips
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
AHASD: Asynchronous Heterogeneous Architecture for LLM Adaptive Drafting Speculative Decoding on Mobile Devices
von: Zirui, Ma, et al.
Veröffentlicht: (2026)
von: Zirui, Ma, et al.
Veröffentlicht: (2026)
SafeCiM: Investigating Resilience of Hybrid Floating-Point Compute-in-Memory Deep Learning Accelerators
von: Bhattacharya, Swastik, et al.
Veröffentlicht: (2025)
von: Bhattacharya, Swastik, et al.
Veröffentlicht: (2025)
SwiftKV: An Edge-Oriented Attention Algorithm and Multi-Head Accelerator for Fast, Efficient LLM Decoding
von: Zhang, Junming, et al.
Veröffentlicht: (2026)
von: Zhang, Junming, et al.
Veröffentlicht: (2026)
Scaling Laws for Floating Point Quantization Training
von: Sun, Xingwu, et al.
Veröffentlicht: (2025)
von: Sun, Xingwu, et al.
Veröffentlicht: (2025)
LLM-DSE: Searching Accelerator Parameters with LLM Agents
von: Wang, Hanyu, et al.
Veröffentlicht: (2025)
von: Wang, Hanyu, et al.
Veröffentlicht: (2025)
E2AFS: Energy-Efficient Approximate Floating Point Square Rooter for Error Tolerant Computing
von: Goyal, Prateek, et al.
Veröffentlicht: (2026)
von: Goyal, Prateek, et al.
Veröffentlicht: (2026)
MXDOTP: A RISC-V ISA Extension for Enabling Microscaling (MX) Floating-Point Dot Products
von: İslamoğlu, Gamze, et al.
Veröffentlicht: (2025)
von: İslamoğlu, Gamze, et al.
Veröffentlicht: (2025)
FuseFPS: Accelerating Farthest Point Sampling with Fusing KD-tree Construction for Point Clouds
von: Han, Meng, et al.
Veröffentlicht: (2023)
von: Han, Meng, et al.
Veröffentlicht: (2023)
Dual-Issue Execution of Mixed Integer and Floating-Point Workloads on Energy-Efficient In-Order RISC-V Cores
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
Exploring and Exploiting Runtime Reconfigurable Floating Point Precision in Scientific Computing: a Case Study for Solving PDEs
von: Hao, Cong "Callie"
Veröffentlicht: (2024)
von: Hao, Cong "Callie"
Veröffentlicht: (2024)
From Natural Language to Silicon: The Representation Bottleneck in LLM Hardware Design
von: Fu, Weimin, et al.
Veröffentlicht: (2026)
von: Fu, Weimin, et al.
Veröffentlicht: (2026)
TransDot: An Area-efficient Reconfigurable Floating-Point Unit for Trans-Precision Dot-Product Accumulation for FPGA AI Engines
von: Wang, Jiayi, et al.
Veröffentlicht: (2026)
von: Wang, Jiayi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness
von: Wang, Huizheng, et al.
Veröffentlicht: (2025) -
SOFA: A Compute-Memory Optimized Sparsity Accelerator via Cross-Stage Coordinated Tiling
von: Wang, Huizheng, et al.
Veröffentlicht: (2024) -
BBAL: A Bidirectional Block Floating Point-Based Quantisation Accelerator for Large Language Models
von: Han, Xiaomeng, et al.
Veröffentlicht: (2025) -
SpecMamba: Accelerating Mamba Inference on FPGA with Speculative Decoding
von: Zhong, Linfeng, et al.
Veröffentlicht: (2025) -
Designing Spatial Architectures for Sparse Attention: STAR Accelerator via Cross-Stage Tiling
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)