FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xia, Haojun, Zheng, Zhen, Wu, Xiaoxia, Chen, Shiyang, Yao, Zhewei, Youn, Stephen, Bakhtiari, Arash, Wyatt, Michael, Zhuang, Donglin, Zhou, Zhongzhu, Ruwase, Olatunji, He, Yuxiong, Song, Shuaiwen Leon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MixFP4: Enhancing NVFP4 with Adaptive FP4/INT4 Block Representations
von: Zou, Jiaxiang, et al.
Veröffentlicht: (2026)
von: Zou, Jiaxiang, et al.
Veröffentlicht: (2026)
DGEMM without FP64 Arithmetic - Using FP64 Emulation and FP8 Tensor Cores with Ozaki Scheme
von: Mukunoki, Daichi
Veröffentlicht: (2025)
von: Mukunoki, Daichi
Veröffentlicht: (2025)
Faster Inference of LLMs using FP8 on the Intel Gaudi
von: Lee, Joonhyung, et al.
Veröffentlicht: (2025)
von: Lee, Joonhyung, et al.
Veröffentlicht: (2025)
Chiplet Cloud: Building AI Supercomputers for Serving Large Generative Language Models
von: Peng, Huwan, et al.
Veröffentlicht: (2023)
von: Peng, Huwan, et al.
Veröffentlicht: (2023)
Hardware-Efficient CNNs: Interleaved Approximate FP32 Multipliers for Kernel Computation
von: Gowda, Bindu G, et al.
Veröffentlicht: (2025)
von: Gowda, Bindu G, et al.
Veröffentlicht: (2025)
FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables
von: Park, Gunho, et al.
Veröffentlicht: (2025)
von: Park, Gunho, et al.
Veröffentlicht: (2025)
Execution-Centric Characterization of FP8 Matrix Cores, Asynchronous Execution, and Structured Sparsity on AMD MI300A
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)
Schrödinger's FP: Dynamic Adaptation of Floating-Point Containers for Deep Learning Training
von: Nikolić, Miloš, et al.
Veröffentlicht: (2022)
von: Nikolić, Miloš, et al.
Veröffentlicht: (2022)
TMA-Adaptive FP8 Grouped GEMM: Eliminating Padding Requirements in Low-Precision Training and Inference on Hopper
von: Su, Zhongling, et al.
Veröffentlicht: (2025)
von: Su, Zhongling, et al.
Veröffentlicht: (2025)
Balancing FP8 Computation Accuracy and Efficiency on Digital CIM via Shift-Aware On-the-fly Aligned-Mantissa Bitwidth Prediction
von: Zhao, Liang, et al.
Veröffentlicht: (2026)
von: Zhao, Liang, et al.
Veröffentlicht: (2026)
Diagnosing FP4 inference: a layer-wise and block-wise sensitivity analysis of NVFP4 and MXFP4
von: Cim, Musa, et al.
Veröffentlicht: (2026)
von: Cim, Musa, et al.
Veröffentlicht: (2026)
LLM-FP4: 4-Bit Floating-Point Quantized Transformers
von: Liu, Shih-yang, et al.
Veröffentlicht: (2023)
von: Liu, Shih-yang, et al.
Veröffentlicht: (2023)
SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
PASCAL: A Phase-Aware Scheduling Algorithm for Serving Reasoning-based Large Language Models
von: Cho, Eunyeong, et al.
Veröffentlicht: (2026)
von: Cho, Eunyeong, et al.
Veröffentlicht: (2026)
Swapping-Centric Neural Recording Systems
von: Ugur, Muhammed, et al.
Veröffentlicht: (2024)
von: Ugur, Muhammed, et al.
Veröffentlicht: (2024)
Adaptive Prognostic Malfunction Based Processor for Autonomous Landing Guidance Assistance System Using FPGA
von: Ahmed, Hossam O., et al.
Veröffentlicht: (2024)
von: Ahmed, Hossam O., et al.
Veröffentlicht: (2024)
ADOR: A Design Exploration Framework for LLM Serving with Enhanced Latency and Throughput
von: Kim, Junsoo, et al.
Veröffentlicht: (2025)
von: Kim, Junsoo, et al.
Veröffentlicht: (2025)
Hardware-Software Co-design for 3D-DRAM-based LLM Serving Accelerator
von: Li, Cong, et al.
Veröffentlicht: (2026)
von: Li, Cong, et al.
Veröffentlicht: (2026)
Mapping Space Exploration for Multi-Chiplet Accelerators Targeting LLM Inference Serving Workloads
von: Li, Boyu, et al.
Veröffentlicht: (2025)
von: Li, Boyu, et al.
Veröffentlicht: (2025)
FastPersist: Accelerating Model Checkpointing in Deep Learning
von: Wang, Guanhua, et al.
Veröffentlicht: (2024)
von: Wang, Guanhua, et al.
Veröffentlicht: (2024)
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
von: Zhou, Zhongzhu, et al.
Veröffentlicht: (2026)
von: Zhou, Zhongzhu, et al.
Veröffentlicht: (2026)
PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
von: Yüzügüler, Ahmet Caner, et al.
Veröffentlicht: (2025)
von: Yüzügüler, Ahmet Caner, et al.
Veröffentlicht: (2025)
Hardware-Centric Analysis of DeepSeek's Multi-Head Latent Attention
von: Geens, Robin, et al.
Veröffentlicht: (2025)
von: Geens, Robin, et al.
Veröffentlicht: (2025)
A 95.5Gb/s 29.6ns worst-case latency ORBGRAND decoder for 6G xURLLC
von: Condo, Carlo
Veröffentlicht: (2024)
von: Condo, Carlo
Veröffentlicht: (2024)
A Switch-Centric In-Network Architecture for Accelerating LLM Inference in Shared-Memory Network
von: Jiang, Aojie, et al.
Veröffentlicht: (2026)
von: Jiang, Aojie, et al.
Veröffentlicht: (2026)
TENET: An Efficient Sparsity-Aware LUT-Centric Architecture for Ternary LLM Inference On Edge
von: Huang, Zhirui, et al.
Veröffentlicht: (2025)
von: Huang, Zhirui, et al.
Veröffentlicht: (2025)
NeuroBlend: Towards Low-Power yet Accurate Neural Network-Based Inference Engine Blending Binary and Fixed-Point Convolutions
von: Fayyazi, Arash, et al.
Veröffentlicht: (2023)
von: Fayyazi, Arash, et al.
Veröffentlicht: (2023)
Nanoscaling Floating-Point (NxFP): NanoMantissa, Adaptive Microexponents, and Code Recycling for Direct-Cast Compression of Large Language Models
von: Lo, Yun-Chen, et al.
Veröffentlicht: (2024)
von: Lo, Yun-Chen, et al.
Veröffentlicht: (2024)
NVLLM: A 3D NAND-Centric Architecture Enabling Edge on-Device LLM Inference
von: Hao, Mingbo, et al.
Veröffentlicht: (2026)
von: Hao, Mingbo, et al.
Veröffentlicht: (2026)
Recurrent CircuitSAT Sampling for Sequential Circuits
von: Ardakani, Arash, et al.
Veröffentlicht: (2025)
von: Ardakani, Arash, et al.
Veröffentlicht: (2025)
Using a Performance Model to Implement a Superscalar CVA6
von: Allart, Côme, et al.
Veröffentlicht: (2024)
von: Allart, Côme, et al.
Veröffentlicht: (2024)
Efficient Trace for RISC-V: Design, Evaluation, and Integration in CVA6
von: Laghi, Umberto, et al.
Veröffentlicht: (2025)
von: Laghi, Umberto, et al.
Veröffentlicht: (2025)
Fixed and Movable Antenna Technology for 6G Integrated Sensing and Communication
von: Zeng, Yong, et al.
Veröffentlicht: (2024)
von: Zeng, Yong, et al.
Veröffentlicht: (2024)
Generalized Ping-Pong: Off-Chip Memory Bandwidth Centric Pipelining Strategy for Processing-In-Memory Accelerators
von: Wang, Ruibao, et al.
Veröffentlicht: (2024)
von: Wang, Ruibao, et al.
Veröffentlicht: (2024)
Development of High-Performance DSP Algorithms on the European Rad-Hard NG-ULTRA SoC FPGA
von: Leon, Vasileios, et al.
Veröffentlicht: (2024)
von: Leon, Vasileios, et al.
Veröffentlicht: (2024)
A 65-nm Reliable 6T CMOS SRAM Cell with Minimum Size Transistors
von: Torrens, Gabriel, et al.
Veröffentlicht: (2024)
von: Torrens, Gabriel, et al.
Veröffentlicht: (2024)
CVA6S+: A Superscalar RISC-V Core with High-Throughput Memory Architecture
von: Tedeschi, Riccardo, et al.
Veröffentlicht: (2025)
von: Tedeschi, Riccardo, et al.
Veröffentlicht: (2025)
From Circuits to SoC Processors: Arithmetic Approximation Techniques & Embedded Computing Methodologies for DSP Acceleration
von: Leon, Vasileios
Veröffentlicht: (2023)
von: Leon, Vasileios
Veröffentlicht: (2023)
GenPairX: A Hardware-Algorithm Co-Designed Accelerator for Paired-End Read Mapping
von: Eudine, Julien, et al.
Veröffentlicht: (2026)
von: Eudine, Julien, et al.
Veröffentlicht: (2026)
Design of a 6-bit Threshold Inverter Quantization (TIQ) Flash Analog to Digital Converter (ADC)
von: Sarkar, Noyon Kumar, et al.
Veröffentlicht: (2025)
von: Sarkar, Noyon Kumar, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MixFP4: Enhancing NVFP4 with Adaptive FP4/INT4 Block Representations
von: Zou, Jiaxiang, et al.
Veröffentlicht: (2026) -
DGEMM without FP64 Arithmetic - Using FP64 Emulation and FP8 Tensor Cores with Ozaki Scheme
von: Mukunoki, Daichi
Veröffentlicht: (2025) -
Faster Inference of LLMs using FP8 on the Intel Gaudi
von: Lee, Joonhyung, et al.
Veröffentlicht: (2025) -
Chiplet Cloud: Building AI Supercomputers for Serving Large Generative Language Models
von: Peng, Huwan, et al.
Veröffentlicht: (2023) -
Hardware-Efficient CNNs: Interleaved Approximate FP32 Multipliers for Kernel Computation
von: Gowda, Bindu G, et al.
Veröffentlicht: (2025)