LLM-FP4: 4-Bit Floating-Point Quantized Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Shih-yang, Liu, Zechun, Huang, Xijie, Dong, Pingcheng, Cheng, Kwang-Ting |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Genetic Quantization-Aware Approximation for Non-Linear Operations in Transformers
by: Dong, Pingcheng, et al.
Published: (2024)
by: Dong, Pingcheng, et al.
Published: (2024)
Scaling Laws for Floating Point Quantization Training
by: Sun, Xingwu, et al.
Published: (2025)
by: Sun, Xingwu, et al.
Published: (2025)
APSQ: Additive Partial Sum Quantization with Algorithm-Hardware Co-Design
by: Tan, Yonghao, et al.
Published: (2025)
by: Tan, Yonghao, et al.
Published: (2025)
Quantization Variation: A New Perspective on Training Transformers with Low-Bit Precision
by: Huang, Xijie, et al.
Published: (2023)
by: Huang, Xijie, et al.
Published: (2023)
SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
Procrastination Is All You Need: Exponent Indexed Accumulators for Floating Point, Posits and Logarithmic Numbers
by: Liguori, Vincenzo
Published: (2024)
by: Liguori, Vincenzo
Published: (2024)
MVQ:Towards Efficient DNN Compression and Acceleration with Masked Vector Quantization
by: Li, Shuaiting, et al.
Published: (2024)
by: Li, Shuaiting, et al.
Published: (2024)
Balancing FP8 Computation Accuracy and Efficiency on Digital CIM via Shift-Aware On-the-fly Aligned-Mantissa Bitwidth Prediction
by: Zhao, Liang, et al.
Published: (2026)
by: Zhao, Liang, et al.
Published: (2026)
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
by: Du, Dayou, et al.
Published: (2025)
by: Du, Dayou, et al.
Published: (2025)
Shedding the Bits: Pushing the Boundaries of Quantization with Minifloats on FPGAs
by: Aggarwal, Shivam, et al.
Published: (2023)
by: Aggarwal, Shivam, et al.
Published: (2023)
Vision Transformer Computation and Resilience for Dynamic Inference
by: Sreedhar, Kavya, et al.
Published: (2022)
by: Sreedhar, Kavya, et al.
Published: (2022)
Efficient and Robust Quantization-aware Training via Adaptive Coreset Selection
by: Huang, Xijie, et al.
Published: (2023)
by: Huang, Xijie, et al.
Published: (2023)
FG-Attn: Leveraging Fine-Grained Sparsity In Diffusion Transformers
by: Durvasula, Sankeerth, et al.
Published: (2025)
by: Durvasula, Sankeerth, et al.
Published: (2025)
Evolving Layer-Specific Scalar Functions for Hardware-Aware Transformer Adaptation
by: Carrigg, Kieran, et al.
Published: (2026)
by: Carrigg, Kieran, et al.
Published: (2026)
Vision Transformers on the Edge: A Comprehensive Survey of Model Compression and Acceleration Strategies
by: Saha, Shaibal, et al.
Published: (2025)
by: Saha, Shaibal, et al.
Published: (2025)
31.1 A 14.08-to-135.69Token/s ReRAM-on-Logic Stacked Outlier-Free Large-Language-Model Accelerator with Block-Clustered Weight-Compression and Adaptive Parallel-Speculative-Decoding
by: Dong, Pingcheng, et al.
Published: (2026)
by: Dong, Pingcheng, et al.
Published: (2026)
CAMO: Correlation-Aware Mask Optimization with Modulated Reinforcement Learning
by: Liang, Xiaoxiao, et al.
Published: (2024)
by: Liang, Xiaoxiao, et al.
Published: (2024)
Fine-Tuning Small Language Models for Domain-Specific AI: An Edge AI Perspective
by: Aralimatti, Rakshit, et al.
Published: (2025)
by: Aralimatti, Rakshit, et al.
Published: (2025)
Understanding and Mitigating Errors of LLM-Generated RTL Code
by: Zhang, Jiazheng, et al.
Published: (2025)
by: Zhang, Jiazheng, et al.
Published: (2025)
Using GUI Agent for Electronic Design Automation
by: Li, Chunyi, et al.
Published: (2025)
by: Li, Chunyi, et al.
Published: (2025)
Efficient stereo matching on embedded GPUs with zero-means cross correlation
by: Chang, Qiong, et al.
Published: (2022)
by: Chang, Qiong, et al.
Published: (2022)
AHCQ-SAM: Toward Accurate and Hardware-Compatible Post-Training Segment Anything Model Quantization
by: Zhang, Wenlun, et al.
Published: (2025)
by: Zhang, Wenlun, et al.
Published: (2025)
InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference
by: Pan, Xiurui, et al.
Published: (2024)
by: Pan, Xiurui, et al.
Published: (2024)
HaLoRA: Hardware-aware Low-Rank Adaptation for Large Language Models Based on Hybrid Compute-in-Memory Architecture
by: Wu, Taiqiang, et al.
Published: (2025)
by: Wu, Taiqiang, et al.
Published: (2025)
Real-Time Spacecraft Pose Estimation Using Mixed-Precision Quantized Neural Network on COTS Reconfigurable MPSoC
by: Posso, Julien, et al.
Published: (2024)
by: Posso, Julien, et al.
Published: (2024)
GRPO with State Mutations: Improving LLM-Based Hardware Test Plan Generation
by: Kochar, Dimple Vijay, et al.
Published: (2026)
by: Kochar, Dimple Vijay, et al.
Published: (2026)
AppSign: Multi-level Approximate Computing for Real-Time Traffic Sign Recognition in Autonomous Vehicles
by: Omidian, Fatemeh, et al.
Published: (2024)
by: Omidian, Fatemeh, et al.
Published: (2024)
Co-designing a Sub-millisecond Latency Event-based Eye Tracking System with Submanifold Sparse CNN
by: Zhang, Baoheng, et al.
Published: (2024)
by: Zhang, Baoheng, et al.
Published: (2024)
ORBIS: Output-Guided Token Reduction with Distribution-Aware Matching for Video Diffusion Acceleration
by: Lee, Hangyeol, et al.
Published: (2026)
by: Lee, Hangyeol, et al.
Published: (2026)
Primitive-Driven Acceleration of Hyperdimensional Computing for Real-Time Image Classification
by: Parikh, Dhruv, et al.
Published: (2026)
by: Parikh, Dhruv, et al.
Published: (2026)
TIMERIPPLE: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Space
by: Miao, Wenxuan, et al.
Published: (2025)
by: Miao, Wenxuan, et al.
Published: (2025)
Identifying Unnecessary 3D Gaussians using Clustering for Fast Rendering of 3D Gaussian Splatting
by: Jo, Joongho, et al.
Published: (2024)
by: Jo, Joongho, et al.
Published: (2024)
Real-Time Object Detection and Classification using YOLO for Edge FPGAs
by: Amin, Rashed Al, et al.
Published: (2025)
by: Amin, Rashed Al, et al.
Published: (2025)
Gen-NeRF: Efficient and Generalizable Neural Radiance Fields via Algorithm-Hardware Co-Design
by: Fu, Yonggan, et al.
Published: (2023)
by: Fu, Yonggan, et al.
Published: (2023)
SteROI-D: System Design and Mapping for Stereo Depth Inference on Regions of Interest
by: Erhardt, Jack, et al.
Published: (2025)
by: Erhardt, Jack, et al.
Published: (2025)
GS-TG: 3D Gaussian Splatting Accelerator with Tile Grouping for Reducing Redundant Sorting while Preserving Rasterization Efficiency
by: Jo, Joongho, et al.
Published: (2025)
by: Jo, Joongho, et al.
Published: (2025)
Neo: Real-Time On-Device 3D Gaussian Splatting with Reuse-and-Update Sorting Acceleration
by: Oh, Changhun, et al.
Published: (2025)
by: Oh, Changhun, et al.
Published: (2025)
A Parameterizable Convolution Accelerator for Embedded Deep Learning Applications
by: Mousouliotis, Panagiotis, et al.
Published: (2026)
by: Mousouliotis, Panagiotis, et al.
Published: (2026)
SpNeRF: Memory Efficient Sparse Volumetric Neural Rendering Accelerator for Edge Devices
by: Zhang, Yipu, et al.
Published: (2025)
by: Zhang, Yipu, et al.
Published: (2025)
Dedicated Inference Engine and Binary-Weight Neural Networks for Lightweight Instance Segmentation
by: Chen, Tse-Wei, et al.
Published: (2025)
by: Chen, Tse-Wei, et al.
Published: (2025)
Similar Items
-
Genetic Quantization-Aware Approximation for Non-Linear Operations in Transformers
by: Dong, Pingcheng, et al.
Published: (2024) -
Scaling Laws for Floating Point Quantization Training
by: Sun, Xingwu, et al.
Published: (2025) -
APSQ: Additive Partial Sum Quantization with Algorithm-Hardware Co-Design
by: Tan, Yonghao, et al.
Published: (2025) -
Quantization Variation: A New Perspective on Training Transformers with Low-Bit Precision
by: Huang, Xijie, et al.
Published: (2023) -
SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
by: Zhang, Jintao, et al.
Published: (2025)