FLIQS: One-Shot Mixed-Precision Floating-Point and Integer Quantization Search
Fuente:
arXiv
Saved in:
| Main Authors: | Dotzel, Jordan, Wu, Gang, Li, Andrew, Umar, Muhammad, Ni, Yun, Abdelfattah, Mohamed S., Zhang, Zhiru, Cheng, Liqun, Dixon, Martin G., Jouppi, Norman P., Le, Quoc V., Li, Sheng |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring the Limits of Semantic Image Compression at Micro-bits per Pixel
by: Dotzel, Jordan, et al.
Published: (2024)
by: Dotzel, Jordan, et al.
Published: (2024)
Semantic Compression of 3D Objects for Open and Collaborative Virtual Worlds
by: Dotzel, Jordan, et al.
Published: (2025)
by: Dotzel, Jordan, et al.
Published: (2025)
Learning from Students: Applying t-Distributions to Explore Accurate and Efficient Formats for LLMs
by: Dotzel, Jordan, et al.
Published: (2024)
by: Dotzel, Jordan, et al.
Published: (2024)
Radial Networks: Dynamic Layer Routing for High-Performance Large Language Models
by: Dotzel, Jordan, et al.
Published: (2024)
by: Dotzel, Jordan, et al.
Published: (2024)
ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models
by: Akhauri, Yash, et al.
Published: (2024)
by: Akhauri, Yash, et al.
Published: (2024)
Analysis of Floating-Point Matrix Multiplication Computed via Integer Arithmetic
by: Abdelfattah, Ahmad, et al.
Published: (2025)
by: Abdelfattah, Ahmad, et al.
Published: (2025)
Kratos: An FPGA Benchmark for Unrolled DNNs with Fine-Grained Sparsity and Mixed Precision
by: Dai, Xilai, et al.
Published: (2024)
by: Dai, Xilai, et al.
Published: (2024)
decoupleQ: Towards 2-bit Post-Training Uniform Quantization via decoupling Parameters into Integer and Floating Points
by: Guo, Yi, et al.
Published: (2024)
by: Guo, Yi, et al.
Published: (2024)
Scaling Laws for Floating Point Quantization Training
by: Sun, Xingwu, et al.
Published: (2025)
by: Sun, Xingwu, et al.
Published: (2025)
Mixed-Precision Quantization for Deep Vision Models with Integer Quadratic Programming
by: Deng, Zihao, et al.
Published: (2023)
by: Deng, Zihao, et al.
Published: (2023)
Range, Not Precision: Block-Floating-Point Half-Precision FFT and SAR Imaging on Apple Silicon
by: Bergach, Mohamed Amine
Published: (2026)
by: Bergach, Mohamed Amine
Published: (2026)
On Latency Predictors for Neural Architecture Search
by: Akhauri, Yash, et al.
Published: (2024)
by: Akhauri, Yash, et al.
Published: (2024)
Mixed-Precision in High-Order Methods: the Impact of Floating-Point Precision on the ADER-DG Algorithm
by: Marot-Lassauzaie, Marc, et al.
Published: (2025)
by: Marot-Lassauzaie, Marc, et al.
Published: (2025)
Encodings for Prediction-based Neural Architecture Search
by: Akhauri, Yash, et al.
Published: (2024)
by: Akhauri, Yash, et al.
Published: (2024)
Dual-Issue Execution of Mixed Integer and Floating-Point Workloads on Energy-Efficient In-Order RISC-V Cores
by: Colagrande, Luca, et al.
Published: (2025)
by: Colagrande, Luca, et al.
Published: (2025)
Direct Search Algorithm for Clock Skew Compensation Immune to Floating-Point Precision Loss
by: Kim, Kyeong Soo
Published: (2025)
by: Kim, Kyeong Soo
Published: (2025)
On Approximate 8-bit Floating-Point Operations Using Integer Operations
by: Lindberg, Theodor, et al.
Published: (2024)
by: Lindberg, Theodor, et al.
Published: (2024)
QuEPT: Quantized Elastic Precision Transformers with One-Shot Calibration for Multi-Bit Switching
by: Xu, Ke, et al.
Published: (2026)
by: Xu, Ke, et al.
Published: (2026)
OMPQ: Orthogonal Mixed Precision Quantization
by: Ma, Yuexiao, et al.
Published: (2021)
by: Ma, Yuexiao, et al.
Published: (2021)
EVCAR-KG: A Knowledge-Infused Multi-Agent Reinforcement Learning Framework for Resilient Electric Vehicle Charging Network Recovery
by: Abdelfattah, Mohamed
Published: (2025)
by: Abdelfattah, Mohamed
Published: (2025)
FPQVAR: Floating Point Quantization for Visual Autoregressive Model with FPGA Hardware Co-design
by: Wei, Renjie, et al.
Published: (2025)
by: Wei, Renjie, et al.
Published: (2025)
Reduced Floating-Point Precision Implicit Monte Carlo
by: Butson, Simon, et al.
Published: (2025)
by: Butson, Simon, et al.
Published: (2025)
Precision Arithmetic: A New Floating-Point Arithmetic
by: Wang, Chengpu
Published: (2006)
by: Wang, Chengpu
Published: (2006)
Satire: Computing Rigorous Bounds for Floating-Point Rounding Error in Mixed-Precision Loop-Free Programs
by: Tirpankar, Tanmay, et al.
Published: (2025)
by: Tirpankar, Tanmay, et al.
Published: (2025)
Search Your Block Floating Point Scales!
by: Gupta, Tanmaey, et al.
Published: (2026)
by: Gupta, Tanmaey, et al.
Published: (2026)
Accurate Reduced Floating-Point Precision Implicit Monte Carlo
by: Butson, Simon, et al.
Published: (2025)
by: Butson, Simon, et al.
Published: (2025)
Trainable Fixed-Point Quantization for Deep Learning Acceleration on FPGAs
by: Dai, Dingyi, et al.
Published: (2024)
by: Dai, Dingyi, et al.
Published: (2024)
F-BFQ: Flexible Block Floating-Point Quantization Accelerator for LLMs
by: Haris, Jude, et al.
Published: (2025)
by: Haris, Jude, et al.
Published: (2025)
LLM-FP4: 4-Bit Floating-Point Quantized Transformers
by: Liu, Shih-yang, et al.
Published: (2023)
by: Liu, Shih-yang, et al.
Published: (2023)
Provably Data-driven Lagrangian Relaxation for Mixed Integer Linear Programming
by: Le, Tung Quoc, et al.
Published: (2026)
by: Le, Tung Quoc, et al.
Published: (2026)
SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs
by: AbouElhamayed, Ahmed F., et al.
Published: (2025)
by: AbouElhamayed, Ahmed F., et al.
Published: (2025)
Channel-Wise Mixed-Precision Quantization for Large Language Models
by: Chen, Zihan, et al.
Published: (2024)
by: Chen, Zihan, et al.
Published: (2024)
Benchmarking Post-Training Quantization of Large Language Models under Microscaling Floating Point Formats
by: Zhang, Manyi, et al.
Published: (2026)
by: Zhang, Manyi, et al.
Published: (2026)
Effective and Efficient Mixed Precision Quantization of Speech Foundation Models
by: Xu, Haoning, et al.
Published: (2025)
by: Xu, Haoning, et al.
Published: (2025)
Optimizing Digital Signal Processing with Half-Precision Floating-Point Arithmetic
by: Dr. Elianore Quasar and Dr. Kaia Rykhard
Published: (2019)
by: Dr. Elianore Quasar and Dr. Kaia Rykhard
Published: (2019)
Clock Skew Compensation Algorithm Immune to Floating-Point Precision Loss
by: Kim, Kyeong Soo, et al.
Published: (2021)
by: Kim, Kyeong Soo, et al.
Published: (2021)
Double-Precision Floating-Point Data Visualizations Using Vulkan API
by: Sozen, Nezihe
Published: (2024)
by: Sozen, Nezihe
Published: (2024)
FP4DiT: Towards Effective Floating Point Quantization for Diffusion Transformers
by: Chen, Ruichen, et al.
Published: (2025)
by: Chen, Ruichen, et al.
Published: (2025)
Low-Bitwidth Floating Point Quantization for Efficient High-Quality Diffusion Models
by: Chen, Cheng, et al.
Published: (2024)
by: Chen, Cheng, et al.
Published: (2024)
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
by: Huang, Wei, et al.
Published: (2023)
by: Huang, Wei, et al.
Published: (2023)
Similar Items
-
Exploring the Limits of Semantic Image Compression at Micro-bits per Pixel
by: Dotzel, Jordan, et al.
Published: (2024) -
Semantic Compression of 3D Objects for Open and Collaborative Virtual Worlds
by: Dotzel, Jordan, et al.
Published: (2025) -
Learning from Students: Applying t-Distributions to Explore Accurate and Efficient Formats for LLMs
by: Dotzel, Jordan, et al.
Published: (2024) -
Radial Networks: Dynamic Layer Routing for High-Performance Large Language Models
by: Dotzel, Jordan, et al.
Published: (2024) -
ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models
by: Akhauri, Yash, et al.
Published: (2024)