MGS: Markov Greedy Sums for Accurate Low-Bitwidth Floating-Point Accumulation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Natesh, Vikas, Kung, H. T., Kong, David |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PQS (Prune, Quantize, and Sort): Low-Bitwidth Accumulation of Dot Products in Neural Network Computations
von: Natesh, Vikas, et al.
Veröffentlicht: (2025)
von: Natesh, Vikas, et al.
Veröffentlicht: (2025)
The AetherFloat Family: Block-Scale-Free Quad-Radix Floating-Point Architectures for AI Accelerators
von: Morisaki, Keita
Veröffentlicht: (2026)
von: Morisaki, Keita
Veröffentlicht: (2026)
Search Your Block Floating Point Scales!
von: Gupta, Tanmaey, et al.
Veröffentlicht: (2026)
von: Gupta, Tanmaey, et al.
Veröffentlicht: (2026)
Schrödinger's FP: Dynamic Adaptation of Floating-Point Containers for Deep Learning Training
von: Nikolić, Miloš, et al.
Veröffentlicht: (2022)
von: Nikolić, Miloš, et al.
Veröffentlicht: (2022)
NeuroScalar: A Deep Learning Framework for Fast, Accurate, and In-the-Wild Cycle-Level Performance Prediction
von: Wadle, Shayne, et al.
Veröffentlicht: (2025)
von: Wadle, Shayne, et al.
Veröffentlicht: (2025)
SafeCiM: Investigating Resilience of Hybrid Floating-Point Compute-in-Memory Deep Learning Accelerators
von: Bhattacharya, Swastik, et al.
Veröffentlicht: (2025)
von: Bhattacharya, Swastik, et al.
Veröffentlicht: (2025)
Scaling Laws for Floating Point Quantization Training
von: Sun, Xingwu, et al.
Veröffentlicht: (2025)
von: Sun, Xingwu, et al.
Veröffentlicht: (2025)
DS-CIM: Digital Stochastic Computing-In-Memory Featuring Accurate OR-Accumulation via Sample Region Remapping for Edge AI Models
von: Shao, Kunming, et al.
Veröffentlicht: (2026)
von: Shao, Kunming, et al.
Veröffentlicht: (2026)
Column-wise Quantization of Weights and Partial Sums for Accurate and Efficient Compute-In-Memory Accelerators
von: Kim, Jiyoon, et al.
Veröffentlicht: (2025)
von: Kim, Jiyoon, et al.
Veröffentlicht: (2025)
Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy
von: Xie, Peichen, et al.
Veröffentlicht: (2025)
von: Xie, Peichen, et al.
Veröffentlicht: (2025)
DOMAC: Differentiable Optimization for High-Speed Multipliers and Multiply-Accumulators
von: Xue, Chenhao, et al.
Veröffentlicht: (2025)
von: Xue, Chenhao, et al.
Veröffentlicht: (2025)
Bespoke Approximation of Multiplication-Accumulation and Activation Targeting Printed Multilayer Perceptrons
von: Afentaki, Florentia, et al.
Veröffentlicht: (2023)
von: Afentaki, Florentia, et al.
Veröffentlicht: (2023)
Introducing Instruction-Accurate Simulators for Performance Estimation of Autotuning Workloads
von: Pelke, Rebecca, et al.
Veröffentlicht: (2025)
von: Pelke, Rebecca, et al.
Veröffentlicht: (2025)
F-BFQ: Flexible Block Floating-Point Quantization Accelerator for LLMs
von: Haris, Jude, et al.
Veröffentlicht: (2025)
von: Haris, Jude, et al.
Veröffentlicht: (2025)
PointODE: Lightweight Point Cloud Learning with Neural Ordinary Differential Equations on Edge
von: Sugiura, Keisuke, et al.
Veröffentlicht: (2025)
von: Sugiura, Keisuke, et al.
Veröffentlicht: (2025)
Ascend HiFloat8 Format for Deep Learning
von: Luo, Yuanyong, et al.
Veröffentlicht: (2024)
von: Luo, Yuanyong, et al.
Veröffentlicht: (2024)
HiFloat4 Format for Language Model Inference
von: Luo, Yuanyong, et al.
Veröffentlicht: (2026)
von: Luo, Yuanyong, et al.
Veröffentlicht: (2026)
Compact Yet Highly Accurate Printed Classifiers Using Sequential Support Vector Machine Circuits
von: Sertaridis, Ilias, et al.
Veröffentlicht: (2025)
von: Sertaridis, Ilias, et al.
Veröffentlicht: (2025)
Floating-Point Multiply-Add with Approximate Normalization for Low-Cost Matrix Engines
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2024)
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2024)
Sorted Weight Sectioning for Energy-Efficient Unstructured Sparse DNNs on Compute-in-Memory Crossbars
von: Farias, Matheus, et al.
Veröffentlicht: (2024)
von: Farias, Matheus, et al.
Veröffentlicht: (2024)
Efficient Reprogramming of Memristive Crossbars for DNNs: Weight Sorting and Bit Stucking
von: Farias, Matheus, et al.
Veröffentlicht: (2024)
von: Farias, Matheus, et al.
Veröffentlicht: (2024)
A2Q+: Improving Accumulator-Aware Weight Quantization
von: Colbert, Ian, et al.
Veröffentlicht: (2024)
von: Colbert, Ian, et al.
Veröffentlicht: (2024)
TransDot: An Area-efficient Reconfigurable Floating-Point Unit for Trans-Precision Dot-Product Accumulation for FPGA AI Engines
von: Wang, Jiayi, et al.
Veröffentlicht: (2026)
von: Wang, Jiayi, et al.
Veröffentlicht: (2026)
TimeFloats: Train-in-Memory with Time-Domain Floating-Point Scalar Products
von: Hashem, Maeesha Binte, et al.
Veröffentlicht: (2024)
von: Hashem, Maeesha Binte, et al.
Veröffentlicht: (2024)
Basis Selection: Low-Rank Decomposition of Pretrained Large Language Models for Target Applications
von: Li, Yang, et al.
Veröffentlicht: (2024)
von: Li, Yang, et al.
Veröffentlicht: (2024)
A Stochastic Rounding-Enabled Low-Precision Floating-Point MAC for DNN Training
von: Ali, Sami Ben, et al.
Veröffentlicht: (2024)
von: Ali, Sami Ben, et al.
Veröffentlicht: (2024)
Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML Fusion
von: Nasr-Esfahany, Arash, et al.
Veröffentlicht: (2025)
von: Nasr-Esfahany, Arash, et al.
Veröffentlicht: (2025)
Low-Energy On-Device Personalization for MCUs
von: Huang, Yushan, et al.
Veröffentlicht: (2024)
von: Huang, Yushan, et al.
Veröffentlicht: (2024)
OPAL: Outlier-Preserved Microscaling Quantization Accelerator for Generative Large Language Models
von: Koo, Jahyun, et al.
Veröffentlicht: (2024)
von: Koo, Jahyun, et al.
Veröffentlicht: (2024)
Benchmarking Ultra-Low-Power $μ$NPUs
von: Millar, Josh, et al.
Veröffentlicht: (2025)
von: Millar, Josh, et al.
Veröffentlicht: (2025)
MDM: Manhattan Distance Mapping of DNN Weights for Parasitic-Resistance-Resilient Memristive Crossbars
von: Farias, Matheus, et al.
Veröffentlicht: (2025)
von: Farias, Matheus, et al.
Veröffentlicht: (2025)
Exploration of Low-Power Flexible Stress Monitoring Classifiers for Conformal Wearables
von: Afentaki, Florentia, et al.
Veröffentlicht: (2025)
von: Afentaki, Florentia, et al.
Veröffentlicht: (2025)
Low-Cost FlashAttention with Fused Exponential and Multiplication Hardware Operators
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2025)
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2025)
Low Power Vision Transformer Accelerator with Hardware-Aware Pruning and Optimized Dataflow
von: Hsiung, Ching-Lin, et al.
Veröffentlicht: (2025)
von: Hsiung, Ching-Lin, et al.
Veröffentlicht: (2025)
Efficient FPGA Implementation of Time-Domain Popcount for Low-Complexity Machine Learning
von: Duan, Shengyu, et al.
Veröffentlicht: (2025)
von: Duan, Shengyu, et al.
Veröffentlicht: (2025)
GCN-ABFT: Low-Cost Online Error Checking for Graph Convolutional Networks
von: Peltekis, Christodoulos, et al.
Veröffentlicht: (2024)
von: Peltekis, Christodoulos, et al.
Veröffentlicht: (2024)
Decentor-V: Lightweight ML Training on Low-Power RISC-V Edge Devices
von: Ribeiro, Marcelo, et al.
Veröffentlicht: (2025)
von: Ribeiro, Marcelo, et al.
Veröffentlicht: (2025)
VEXP: A Low-Cost RISC-V ISA Extension for Accelerated Softmax Computation in Transformers
von: Wang, Run, et al.
Veröffentlicht: (2025)
von: Wang, Run, et al.
Veröffentlicht: (2025)
HW-SW Optimization of DNNs for Privacy-preserving People Counting on Low-resolution Infrared Arrays
von: Risso, Matteo, et al.
Veröffentlicht: (2024)
von: Risso, Matteo, et al.
Veröffentlicht: (2024)
Architectural Implications of Neural Network Inference for High Data-Rate, Low-Latency Scientific Applications
von: Weng, Olivia, et al.
Veröffentlicht: (2024)
von: Weng, Olivia, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
PQS (Prune, Quantize, and Sort): Low-Bitwidth Accumulation of Dot Products in Neural Network Computations
von: Natesh, Vikas, et al.
Veröffentlicht: (2025) -
The AetherFloat Family: Block-Scale-Free Quad-Radix Floating-Point Architectures for AI Accelerators
von: Morisaki, Keita
Veröffentlicht: (2026) -
Search Your Block Floating Point Scales!
von: Gupta, Tanmaey, et al.
Veröffentlicht: (2026) -
Schrödinger's FP: Dynamic Adaptation of Floating-Point Containers for Deep Learning Training
von: Nikolić, Miloš, et al.
Veröffentlicht: (2022) -
NeuroScalar: A Deep Learning Framework for Fast, Accurate, and In-the-Wild Cycle-Level Performance Prediction
von: Wadle, Shayne, et al.
Veröffentlicht: (2025)