QERA: an Analytical Framework for Quantization Error Reconstruction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Cheng, Wong, Jeffrey T. H., Xiao, Can, Constantinides, George A., Zhao, Yiren |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LQER: Low-Rank Quantization Error Reconstruction for LLMs
von: Zhang, Cheng, et al.
Veröffentlicht: (2024)
von: Zhang, Cheng, et al.
Veröffentlicht: (2024)
A3 : an Analytical Low-Rank Approximation Framework for Attention
von: Wong, Jeffrey T. H., et al.
Veröffentlicht: (2025)
von: Wong, Jeffrey T. H., et al.
Veröffentlicht: (2025)
Revisiting Block-based Quantisation: What is Important for Sub-8-bit LLM Inference?
von: Zhang, Cheng, et al.
Veröffentlicht: (2023)
von: Zhang, Cheng, et al.
Veröffentlicht: (2023)
AMPLE: Event-Driven Accelerator for Mixed-Precision Inference of Graph Neural Networks
von: Gimenes, Pedro, et al.
Veröffentlicht: (2025)
von: Gimenes, Pedro, et al.
Veröffentlicht: (2025)
Optimised Grouped-Query Attention Mechanism for Transformers
von: Chen, Yuang, et al.
Veröffentlicht: (2024)
von: Chen, Yuang, et al.
Veröffentlicht: (2024)
Unlocking the Global Synergies in Low-Rank Adapters
von: Zhang, Zixi, et al.
Veröffentlicht: (2024)
von: Zhang, Zixi, et al.
Veröffentlicht: (2024)
NeuraLUT-Assemble: Hardware-aware Assembling of Sub-Neural Networks for Efficient LUT Inference
von: Andronic, Marta, et al.
Veröffentlicht: (2025)
von: Andronic, Marta, et al.
Veröffentlicht: (2025)
ARIES: Autonomous Reasoning with LLMs on Interactive Thought Graph Environments
von: Gimenes, Pedro, et al.
Veröffentlicht: (2025)
von: Gimenes, Pedro, et al.
Veröffentlicht: (2025)
On the Existence and Behavior of Secondary Attention Sinks
von: Wong, Jeffrey T. H., et al.
Veröffentlicht: (2025)
von: Wong, Jeffrey T. H., et al.
Veröffentlicht: (2025)
Scaling Laws For Mixed Quantization
von: Cao, Zeyu, et al.
Veröffentlicht: (2024)
von: Cao, Zeyu, et al.
Veröffentlicht: (2024)
NeuraLUT: Hiding Neural Network Density in Boolean Synthesizable Functions
von: Andronic, Marta, et al.
Veröffentlicht: (2024)
von: Andronic, Marta, et al.
Veröffentlicht: (2024)
PolyLUT: Learning Piecewise Polynomials for Ultra-Low Latency FPGA LUT-based Inference
von: Andronic, Marta, et al.
Veröffentlicht: (2023)
von: Andronic, Marta, et al.
Veröffentlicht: (2023)
Quantamination: Dynamic Quantization Leaks Your Data Across the Batch
von: Foerster, Hanna, et al.
Veröffentlicht: (2026)
von: Foerster, Hanna, et al.
Veröffentlicht: (2026)
Preserve-Then-Quantize: Balancing Rank Budgets for Quantization Error Reconstruction in LLMs
von: Cho, Yoonjun, et al.
Veröffentlicht: (2026)
von: Cho, Yoonjun, et al.
Veröffentlicht: (2026)
ASER: Activation Smoothing and Error Reconstruction for Large Language Model Quantization
von: Zhao, Weibo, et al.
Veröffentlicht: (2024)
von: Zhao, Weibo, et al.
Veröffentlicht: (2024)
TriAxialKV: Toward Extreme Low-Precision KV-Cache Quantization for Agentic Inference Tasks
von: Shen, Hanzhang, et al.
Veröffentlicht: (2026)
von: Shen, Hanzhang, et al.
Veröffentlicht: (2026)
Direction-Preserving Number Representations
von: Zadeh, Bardia, et al.
Veröffentlicht: (2026)
von: Zadeh, Bardia, et al.
Veröffentlicht: (2026)
SERQ: Saliency-Aware Low-Rank Error Reconstruction for LLM Quantization
von: Park, Yeonsik, et al.
Veröffentlicht: (2026)
von: Park, Yeonsik, et al.
Veröffentlicht: (2026)
PolyLUT: Ultra-low Latency Polynomial Inference with Hardware-Aware Structured Pruning
von: Andronic, Marta, et al.
Veröffentlicht: (2025)
von: Andronic, Marta, et al.
Veröffentlicht: (2025)
Convergence for Discrete Parameter Update Schemes
von: Wilson, Paul, et al.
Veröffentlicht: (2025)
von: Wilson, Paul, et al.
Veröffentlicht: (2025)
FAQ: Mitigating Quantization Error via Regenerating Calibration Data with Family-Aware Quantization
von: Xiao, Haiyang, et al.
Veröffentlicht: (2026)
von: Xiao, Haiyang, et al.
Veröffentlicht: (2026)
ATHEENA: A Toolflow for Hardware Early-Exit Network Automation
von: Biggs, Benjamin, et al.
Veröffentlicht: (2023)
von: Biggs, Benjamin, et al.
Veröffentlicht: (2023)
Exploring FPGA designs for MX and beyond
von: Samson, Ebby, et al.
Veröffentlicht: (2024)
von: Samson, Ebby, et al.
Veröffentlicht: (2024)
ReducedLUT: Table Decomposition with "Don't Care" Conditions
von: Cassidy, Oliver, et al.
Veröffentlicht: (2024)
von: Cassidy, Oliver, et al.
Veröffentlicht: (2024)
Training with Fewer Bits: Unlocking Edge LLMs Training with Stochastic Rounding
von: Liu, Taowen, et al.
Veröffentlicht: (2025)
von: Liu, Taowen, et al.
Veröffentlicht: (2025)
Hardware and Software Platform Inference
von: Zhang, Cheng, et al.
Veröffentlicht: (2024)
von: Zhang, Cheng, et al.
Veröffentlicht: (2024)
Deep Kernel Fusion for Transformers
von: Zhang, Zixi, et al.
Veröffentlicht: (2026)
von: Zhang, Zixi, et al.
Veröffentlicht: (2026)
Rethinking Residual Errors in Compensation-based LLM Quantization
von: Li, Shuaiting, et al.
Veröffentlicht: (2026)
von: Li, Shuaiting, et al.
Veröffentlicht: (2026)
The Fourth State: Signed-Zero Ternary for Stable LLM Quantization (and More)
von: Uhlmann, Jeffrey
Veröffentlicht: (2025)
von: Uhlmann, Jeffrey
Veröffentlicht: (2025)
Quantization Error Propagation: Revisiting Layer-Wise Post-Training Quantization
von: Arai, Yamato, et al.
Veröffentlicht: (2025)
von: Arai, Yamato, et al.
Veröffentlicht: (2025)
SOAR: Scale Optimization for Accurate Reconstruction in NVFP4 Quantization
von: Bao, Chengzhu, et al.
Veröffentlicht: (2026)
von: Bao, Chengzhu, et al.
Veröffentlicht: (2026)
Predicting Probabilities of Error to Combine Quantization and Early Exiting: QuEE
von: Regol, Florence, et al.
Veröffentlicht: (2024)
von: Regol, Florence, et al.
Veröffentlicht: (2024)
The Exploration of Error Bounds in Classification with Noisy Labels
von: Liu, Haixia, et al.
Veröffentlicht: (2025)
von: Liu, Haixia, et al.
Veröffentlicht: (2025)
Dissecting Quantization Error: A Concentration-Alignment Perspective
von: Federici, Marco, et al.
Veröffentlicht: (2026)
von: Federici, Marco, et al.
Veröffentlicht: (2026)
Pseudo-Quantized Actor-Critic Algorithm for Robustness to Noisy Temporal Difference Error
von: Kobayashi, Taisuke
Veröffentlicht: (2026)
von: Kobayashi, Taisuke
Veröffentlicht: (2026)
OJBKQ: Objective-Joint Babai-Klein Quantization
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
BiSup: Bidirectional Quantization Error Suppression for Large Language Models
von: Zou, Minghui, et al.
Veröffentlicht: (2024)
von: Zou, Minghui, et al.
Veröffentlicht: (2024)
Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction
von: Chhugani, Jatin, et al.
Veröffentlicht: (2026)
von: Chhugani, Jatin, et al.
Veröffentlicht: (2026)
Runtime-Certified Bounded-Error Quantized Attention
von: Calver, Dean
Veröffentlicht: (2026)
von: Calver, Dean
Veröffentlicht: (2026)
Error Diffusion: Post Training Quantization with Block-Scaled Number Formats for Neural Networks
von: Khodamoradi, Alireza, et al.
Veröffentlicht: (2024)
von: Khodamoradi, Alireza, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LQER: Low-Rank Quantization Error Reconstruction for LLMs
von: Zhang, Cheng, et al.
Veröffentlicht: (2024) -
A3 : an Analytical Low-Rank Approximation Framework for Attention
von: Wong, Jeffrey T. H., et al.
Veröffentlicht: (2025) -
Revisiting Block-based Quantisation: What is Important for Sub-8-bit LLM Inference?
von: Zhang, Cheng, et al.
Veröffentlicht: (2023) -
AMPLE: Event-Driven Accelerator for Mixed-Precision Inference of Graph Neural Networks
von: Gimenes, Pedro, et al.
Veröffentlicht: (2025) -
Optimised Grouped-Query Attention Mechanism for Transformers
von: Chen, Yuang, et al.
Veröffentlicht: (2024)