A2Q+: Improving Accumulator-Aware Weight Quantization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Colbert, Ian, Pappalardo, Alessandro, Petri-Koenig, Jakoba, Umuroglu, Yaman |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SIRA: Scaled-Integer Range Analysis for Optimizing FPGA Dataflow Neural Network Accelerators
von: Umuroglu, Yaman, et al.
Veröffentlicht: (2025)
von: Umuroglu, Yaman, et al.
Veröffentlicht: (2025)
PoTAcc: A Pipeline for End-to-End Acceleration of Power-of-Two Quantized DNNs
von: Saha, Rappy, et al.
Veröffentlicht: (2026)
von: Saha, Rappy, et al.
Veröffentlicht: (2026)
SONIQ: System-Optimized Noise-Injected Ultra-Low-Precision Quantization with Full-Precision Parity
von: Zhou, Cyrus, et al.
Veröffentlicht: (2023)
von: Zhou, Cyrus, et al.
Veröffentlicht: (2023)
Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction
von: Chhugani, Jatin, et al.
Veröffentlicht: (2026)
von: Chhugani, Jatin, et al.
Veröffentlicht: (2026)
Shedding the Bits: Pushing the Boundaries of Quantization with Minifloats on FPGAs
von: Aggarwal, Shivam, et al.
Veröffentlicht: (2023)
von: Aggarwal, Shivam, et al.
Veröffentlicht: (2023)
Toward A Formalized Approach for Spike Sorting Algorithms and Hardware Evaluation
von: Zhang, Tim, et al.
Veröffentlicht: (2022)
von: Zhang, Tim, et al.
Veröffentlicht: (2022)
FINN-GL: Generalized Mixed-Precision Extensions for FPGA-Accelerated LSTMs
von: Khandelwal, Shashwat, et al.
Veröffentlicht: (2025)
von: Khandelwal, Shashwat, et al.
Veröffentlicht: (2025)
USEFUSE: Uniform Stride for Enhanced Performance in Fused Layer Architecture of Deep Neural Networks
von: Ibrahim, Muhammad Sohail, et al.
Veröffentlicht: (2024)
von: Ibrahim, Muhammad Sohail, et al.
Veröffentlicht: (2024)
Understanding the Performance Horizon of the Latest ML Workloads with NonGEMM Workloads
von: Karami, Rachid, et al.
Veröffentlicht: (2024)
von: Karami, Rachid, et al.
Veröffentlicht: (2024)
Graph neural networks with configuration cross-attention for tensor compilers
von: Khizbullin, Dmitrii, et al.
Veröffentlicht: (2024)
von: Khizbullin, Dmitrii, et al.
Veröffentlicht: (2024)
Search Your Block Floating Point Scales!
von: Gupta, Tanmaey, et al.
Veröffentlicht: (2026)
von: Gupta, Tanmaey, et al.
Veröffentlicht: (2026)
GCL-Sampler: Discovering Kernel Similarity for Sampled GPU Simulation via Graph Contrastive Learning
von: Wang, Jiaqi, et al.
Veröffentlicht: (2026)
von: Wang, Jiaqi, et al.
Veröffentlicht: (2026)
Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML Fusion
von: Nasr-Esfahany, Arash, et al.
Veröffentlicht: (2025)
von: Nasr-Esfahany, Arash, et al.
Veröffentlicht: (2025)
Prefill vs. Decode Bottlenecks: SRAM-Frequency Tradeoffs and the Memory-Bandwidth Ceiling
von: Atmer, Hannah, et al.
Veröffentlicht: (2025)
von: Atmer, Hannah, et al.
Veröffentlicht: (2025)
Improving Quantization with Post-Training Model Expansion
von: Franco, Giuseppe, et al.
Veröffentlicht: (2025)
von: Franco, Giuseppe, et al.
Veröffentlicht: (2025)
Design Space Exploration of Approximate Computing Techniques with a Reinforcement Learning Approach
von: Saeedi, Sepide, et al.
Veröffentlicht: (2023)
von: Saeedi, Sepide, et al.
Veröffentlicht: (2023)
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
von: Zhang, Hang, et al.
Veröffentlicht: (2025)
von: Zhang, Hang, et al.
Veröffentlicht: (2025)
ASPO: Constraint-Aware Bayesian Optimization for FPGA-based Soft Processors
von: Wu, Haoran, et al.
Veröffentlicht: (2025)
von: Wu, Haoran, et al.
Veröffentlicht: (2025)
LowRA: Accurate and Efficient LoRA Fine-Tuning of LLMs under 2 Bits
von: Zhou, Zikai, et al.
Veröffentlicht: (2025)
von: Zhou, Zikai, et al.
Veröffentlicht: (2025)
RooflineBench: A Benchmarking Framework for On-Device LLMs via Roofline Analysis
von: Bi, Zhen, et al.
Veröffentlicht: (2026)
von: Bi, Zhen, et al.
Veröffentlicht: (2026)
FlexiSAGA: A Flexible Systolic Array GEMM Accelerator for Sparse and Dense Processing
von: Müller, Mika Markus, et al.
Veröffentlicht: (2025)
von: Müller, Mika Markus, et al.
Veröffentlicht: (2025)
SAHM: State-Aware Heterogeneous Multicore for Single-Thread Performance
von: Wadle, Shayne, et al.
Veröffentlicht: (2025)
von: Wadle, Shayne, et al.
Veröffentlicht: (2025)
Automatic Generation of Fast and Accurate Performance Models for Deep Neural Network Accelerators
von: Lübeck, Konstantin, et al.
Veröffentlicht: (2024)
von: Lübeck, Konstantin, et al.
Veröffentlicht: (2024)
It's all about PR -- Smart Benchmarking AI Accelerators using Performance Representatives
von: Jung, Alexander Louis-Ferdinand, et al.
Veröffentlicht: (2024)
von: Jung, Alexander Louis-Ferdinand, et al.
Veröffentlicht: (2024)
Characterizing and Understanding HGNN Training on GPUs
von: Han, Dengke, et al.
Veröffentlicht: (2024)
von: Han, Dengke, et al.
Veröffentlicht: (2024)
NSFlow: An End-to-End FPGA Framework with Scalable Dataflow Architecture for Neuro-Symbolic AI
von: Yang, Hanchen, et al.
Veröffentlicht: (2025)
von: Yang, Hanchen, et al.
Veröffentlicht: (2025)
DEEP-GAP: Deep-learning Evaluation of Execution Parallelism in GPU Architectural Performance
von: Palaniappan, Kathiravan
Veröffentlicht: (2026)
von: Palaniappan, Kathiravan
Veröffentlicht: (2026)
Forecasting LLM Inference Performance via Hardware-Agnostic Analytical Modeling
von: Patwari, Rajeev, et al.
Veröffentlicht: (2025)
von: Patwari, Rajeev, et al.
Veröffentlicht: (2025)
Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes
von: Hendria, Willy Fitra
Veröffentlicht: (2026)
von: Hendria, Willy Fitra
Veröffentlicht: (2026)
SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs
von: AbouElhamayed, Ahmed F., et al.
Veröffentlicht: (2025)
von: AbouElhamayed, Ahmed F., et al.
Veröffentlicht: (2025)
Confidential LLM Inference: Performance and Cost Across CPU and GPU TEEs
von: Chrapek, Marcin, et al.
Veröffentlicht: (2025)
von: Chrapek, Marcin, et al.
Veröffentlicht: (2025)
Adaptive Cache Pollution Control for Large Language Model Inference Workloads Using Temporal CNN-Based Prediction and Priority-Aware Replacement
von: Liu, Songze, et al.
Veröffentlicht: (2025)
von: Liu, Songze, et al.
Veröffentlicht: (2025)
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
von: Fan, Ruibo, et al.
Veröffentlicht: (2026)
von: Fan, Ruibo, et al.
Veröffentlicht: (2026)
Model Quantization and Hardware Acceleration for Vision Transformers: A Comprehensive Survey
von: Du, Dayou, et al.
Veröffentlicht: (2024)
von: Du, Dayou, et al.
Veröffentlicht: (2024)
Systematic Characterization of LLM Quantization: A Performance, Energy, and Quality Perspective
von: Shi, Tianyao, et al.
Veröffentlicht: (2025)
von: Shi, Tianyao, et al.
Veröffentlicht: (2025)
Toward Capturing Genetic Epistasis From Multivariate Genome-Wide Association Studies Using Mixed-Precision Kernel Ridge Regression
von: Ltaief, Hatem, et al.
Veröffentlicht: (2024)
von: Ltaief, Hatem, et al.
Veröffentlicht: (2024)
D-com: Accelerating Iterative Processing to Enable Low-rank Decomposition of Activations
von: Tahmasebi, Faraz, et al.
Veröffentlicht: (2025)
von: Tahmasebi, Faraz, et al.
Veröffentlicht: (2025)
ETM2: Empowering Traditional Memory Bandwidth Regulation using ETM
von: Zuepke, Alexander, et al.
Veröffentlicht: (2026)
von: Zuepke, Alexander, et al.
Veröffentlicht: (2026)
JSPIM: A Skew-Aware PIM Accelerator for High-Performance Databases Join and Select Operations
von: Tajdari, Sabiha, et al.
Veröffentlicht: (2025)
von: Tajdari, Sabiha, et al.
Veröffentlicht: (2025)
Cleaning up the Mess: Re-Evaluating the Real-System Modeling Accuracy of Ramulator 2.0
von: Bostanci, F. Nisa, et al.
Veröffentlicht: (2025)
von: Bostanci, F. Nisa, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SIRA: Scaled-Integer Range Analysis for Optimizing FPGA Dataflow Neural Network Accelerators
von: Umuroglu, Yaman, et al.
Veröffentlicht: (2025) -
PoTAcc: A Pipeline for End-to-End Acceleration of Power-of-Two Quantized DNNs
von: Saha, Rappy, et al.
Veröffentlicht: (2026) -
SONIQ: System-Optimized Noise-Injected Ultra-Low-Precision Quantization with Full-Precision Parity
von: Zhou, Cyrus, et al.
Veröffentlicht: (2023) -
Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction
von: Chhugani, Jatin, et al.
Veröffentlicht: (2026) -
Shedding the Bits: Pushing the Boundaries of Quantization with Minifloats on FPGAs
von: Aggarwal, Shivam, et al.
Veröffentlicht: (2023)