Scaling Laws for Floating Point Quantization Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sun, Xingwu, Li, Shuaipeng, Xie, Ruobing, Han, Weidong, Wu, Kan, Yang, Zhen, Li, Yixing, Wang, An, Li, Shuai, Xue, Jinbao, Cheng, Yu, Tao, Yangyu, Kang, Zhanhui, Xu, Chengzhong, Wang, Di, Jiang, Jie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TimeFloats: Train-in-Memory with Time-Domain Floating-Point Scalar Products
von: Hashem, Maeesha Binte, et al.
Veröffentlicht: (2024)
von: Hashem, Maeesha Binte, et al.
Veröffentlicht: (2024)
A Hybrid-Domain Floating-Point Compute-in-Memory Architecture for Efficient Acceleration of High-Precision Deep Neural Networks
von: Yi, Zhiqiang, et al.
Veröffentlicht: (2025)
von: Yi, Zhiqiang, et al.
Veröffentlicht: (2025)
A Stochastic Rounding-Enabled Low-Precision Floating-Point MAC for DNN Training
von: Ali, Sami Ben, et al.
Veröffentlicht: (2024)
von: Ali, Sami Ben, et al.
Veröffentlicht: (2024)
TransDot: An Area-efficient Reconfigurable Floating-Point Unit for Trans-Precision Dot-Product Accumulation for FPGA AI Engines
von: Wang, Jiayi, et al.
Veröffentlicht: (2026)
von: Wang, Jiayi, et al.
Veröffentlicht: (2026)
Search Your Block Floating Point Scales!
von: Gupta, Tanmaey, et al.
Veröffentlicht: (2026)
von: Gupta, Tanmaey, et al.
Veröffentlicht: (2026)
Schrödinger's FP: Dynamic Adaptation of Floating-Point Containers for Deep Learning Training
von: Nikolić, Miloš, et al.
Veröffentlicht: (2022)
von: Nikolić, Miloš, et al.
Veröffentlicht: (2022)
BBAL: A Bidirectional Block Floating Point-Based Quantisation Accelerator for Large Language Models
von: Han, Xiaomeng, et al.
Veröffentlicht: (2025)
von: Han, Xiaomeng, et al.
Veröffentlicht: (2025)
The AetherFloat Family: Block-Scale-Free Quad-Radix Floating-Point Architectures for AI Accelerators
von: Morisaki, Keita
Veröffentlicht: (2026)
von: Morisaki, Keita
Veröffentlicht: (2026)
Fast Generation of Custom Floating-Point Spatial Filters on FPGAs
von: Campos, Nelson, et al.
Veröffentlicht: (2024)
von: Campos, Nelson, et al.
Veröffentlicht: (2024)
Online Alignment and Addition in Multi-Term Floating-Point Adders
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2024)
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2024)
Floating Point HUB Adder for RISC-V Sargantana Processor
von: Bandera, Gerardo, et al.
Veröffentlicht: (2024)
von: Bandera, Gerardo, et al.
Veröffentlicht: (2024)
Floating-Point Multiply-Add with Approximate Normalization for Low-Cost Matrix Engines
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2024)
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2024)
Converting Binary Floating-Point Numbers to Shortest Decimal Strings: An Experimental Review
von: Gareau, Jaël Champagne, et al.
Veröffentlicht: (2026)
von: Gareau, Jaël Champagne, et al.
Veröffentlicht: (2026)
From Quarter to All: Accelerating Speculative LLM Decoding via Floating-Point Exponent Remapping and Parameter Sharing
von: Zhao, Yushu, et al.
Veröffentlicht: (2025)
von: Zhao, Yushu, et al.
Veröffentlicht: (2025)
H-FA: A Hybrid Floating-Point and Logarithmic Approach to Hardware Accelerated FlashAttention
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2025)
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2025)
MXFormer: A Microscaling Floating-Point Charge-Trap Transistor Compute-in-Memory Transformer Accelerator
von: Karfakis, George, et al.
Veröffentlicht: (2026)
von: Karfakis, George, et al.
Veröffentlicht: (2026)
E2AFS: Energy-Efficient Approximate Floating Point Square Rooter for Error Tolerant Computing
von: Goyal, Prateek, et al.
Veröffentlicht: (2026)
von: Goyal, Prateek, et al.
Veröffentlicht: (2026)
MXDOTP: A RISC-V ISA Extension for Enabling Microscaling (MX) Floating-Point Dot Products
von: İslamoğlu, Gamze, et al.
Veröffentlicht: (2025)
von: İslamoğlu, Gamze, et al.
Veröffentlicht: (2025)
Unicorn-CIM: Uncovering the Vulnerability and Improving the Resilience of High-Precision Compute-in-Memory
von: Li, Qiufeng, et al.
Veröffentlicht: (2025)
von: Li, Qiufeng, et al.
Veröffentlicht: (2025)
Dual-Issue Execution of Mixed Integer and Floating-Point Workloads on Energy-Efficient In-Order RISC-V Cores
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
Exploring and Exploiting Runtime Reconfigurable Floating Point Precision in Scientific Computing: a Case Study for Solving PDEs
von: Hao, Cong "Callie"
Veröffentlicht: (2024)
von: Hao, Cong "Callie"
Veröffentlicht: (2024)
Inexactness and Correction of Floating-Point Reciprocal, Division and Square Root
von: Dutton, Lucas M., et al.
Veröffentlicht: (2024)
von: Dutton, Lucas M., et al.
Veröffentlicht: (2024)
FuseFPS: Accelerating Farthest Point Sampling with Fusing KD-tree Construction for Point Clouds
von: Han, Meng, et al.
Veröffentlicht: (2023)
von: Han, Meng, et al.
Veröffentlicht: (2023)
Critical Path Aware Timing-Driven Global Placement for Large-Scale Heterogeneous FPGAs
von: Jiang, He, et al.
Veröffentlicht: (2025)
von: Jiang, He, et al.
Veröffentlicht: (2025)
Range, Not Precision: Block-Floating-Point Half-Precision FFT and SAR Imaging on Apple Silicon
von: Bergach, Mohamed Amine
Veröffentlicht: (2026)
von: Bergach, Mohamed Amine
Veröffentlicht: (2026)
F-BFQ: Flexible Block Floating-Point Quantization Accelerator for LLMs
von: Haris, Jude, et al.
Veröffentlicht: (2025)
von: Haris, Jude, et al.
Veröffentlicht: (2025)
On Approximate 8-bit Floating-Point Operations Using Integer Operations
von: Lindberg, Theodor, et al.
Veröffentlicht: (2024)
von: Lindberg, Theodor, et al.
Veröffentlicht: (2024)
Closing the Gap Between Float and Posit Hardware Efficiency
von: Jonnalagadda, Aditya Anirudh, et al.
Veröffentlicht: (2026)
von: Jonnalagadda, Aditya Anirudh, et al.
Veröffentlicht: (2026)
In-Memory ADC-Based Nonlinear Activation Quantization for Efficient In-Memory Computing
von: Dong, Shuai, et al.
Veröffentlicht: (2026)
von: Dong, Shuai, et al.
Veröffentlicht: (2026)
DEFA: Efficient Deformable Attention Acceleration via Pruning-Assisted Grid-Sampling and Multi-Scale Parallel Processing
von: Xu, Yansong, et al.
Veröffentlicht: (2024)
von: Xu, Yansong, et al.
Veröffentlicht: (2024)
MiniFloat-NN and ExSdotp: An ISA Extension and a Modular Open Hardware Unit for Low-Precision Training on RISC-V cores
von: Bertaccini, Luca, et al.
Veröffentlicht: (2022)
von: Bertaccini, Luca, et al.
Veröffentlicht: (2022)
SSRESF: Sensitivity-aware Single-particle Radiation Effects Simulation Framework in SoC Platforms based on SVM Algorithm
von: Liu, Meng, et al.
Veröffentlicht: (2024)
von: Liu, Meng, et al.
Veröffentlicht: (2024)
MGS: Markov Greedy Sums for Accurate Low-Bitwidth Floating-Point Accumulation
von: Natesh, Vikas, et al.
Veröffentlicht: (2025)
von: Natesh, Vikas, et al.
Veröffentlicht: (2025)
Aging Aware Adaptive Voltage Scaling for Reliable and Efficient AI Accelerators
von: Xie, Tong, et al.
Veröffentlicht: (2026)
von: Xie, Tong, et al.
Veröffentlicht: (2026)
A Tensor-Train Decomposition based Compression of LLMs on Group Vector Systolic Accelerator
von: Huang, Sixiao, et al.
Veröffentlicht: (2025)
von: Huang, Sixiao, et al.
Veröffentlicht: (2025)
Mozart: Modularized and Efficient MoE Training on 3.5D Wafer-Scale Chiplet Architectures
von: Luo, Shuqing, et al.
Veröffentlicht: (2026)
von: Luo, Shuqing, et al.
Veröffentlicht: (2026)
Formal that "Floats" High: Formal Verification of Floating Point Arithmetic
von: Mohanty, Hansa, et al.
Veröffentlicht: (2025)
von: Mohanty, Hansa, et al.
Veröffentlicht: (2025)
When Forgetting Builds Reliability: LLM Unlearning for Reliable Hardware Code Generation
von: Liang, Yiwen, et al.
Veröffentlicht: (2025)
von: Liang, Yiwen, et al.
Veröffentlicht: (2025)
Accelerator-assisted Floating-point ASIP for Communication and Positioning in Massive MIMO Systems
von: Attari, Mohammad, et al.
Veröffentlicht: (2025)
von: Attari, Mohammad, et al.
Veröffentlicht: (2025)
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge
von: Tian, Chunlin, et al.
Veröffentlicht: (2025)
von: Tian, Chunlin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
TimeFloats: Train-in-Memory with Time-Domain Floating-Point Scalar Products
von: Hashem, Maeesha Binte, et al.
Veröffentlicht: (2024) -
A Hybrid-Domain Floating-Point Compute-in-Memory Architecture for Efficient Acceleration of High-Precision Deep Neural Networks
von: Yi, Zhiqiang, et al.
Veröffentlicht: (2025) -
A Stochastic Rounding-Enabled Low-Precision Floating-Point MAC for DNN Training
von: Ali, Sami Ben, et al.
Veröffentlicht: (2024) -
TransDot: An Area-efficient Reconfigurable Floating-Point Unit for Trans-Precision Dot-Product Accumulation for FPGA AI Engines
von: Wang, Jiayi, et al.
Veröffentlicht: (2026) -
Search Your Block Floating Point Scales!
von: Gupta, Tanmaey, et al.
Veröffentlicht: (2026)