LO-BCQ: Block Clustered Quantization for 4-bit (W4A4) LLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Elangovan, Reena, Sakr, Charbel, Raghunathan, Anand, Khailany, Brucek |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ESPACE: Dimensionality Reduction of Activations for Model Compression
by: Sakr, Charbel, et al.
Published: (2024)
by: Sakr, Charbel, et al.
Published: (2024)
QuRL: Efficient Reinforcement Learning with Quantized Rollout
by: Li, Yuhang, et al.
Published: (2026)
by: Li, Yuhang, et al.
Published: (2026)
FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference
by: Hooper, Coleman, et al.
Published: (2025)
by: Hooper, Coleman, et al.
Published: (2025)
ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning Models
by: Ramachandran, Akshat, et al.
Published: (2025)
by: Ramachandran, Akshat, et al.
Published: (2025)
LLM4Cov: Execution-Aware Agentic Learning for High-coverage Testbench Generation
by: Zhang, Hejia, et al.
Published: (2026)
by: Zhang, Hejia, et al.
Published: (2026)
Improving Block-Wise LLM Quantization by 4-bit Block-Wise Optimal Float (BOF4): Analysis and Variations
by: Blumenberg, Patrick, et al.
Published: (2025)
by: Blumenberg, Patrick, et al.
Published: (2025)
SQ-DM: Accelerating Diffusion Models with Aggressive Quantization and Temporal Sparsity
by: Fan, Zichen, et al.
Published: (2025)
by: Fan, Zichen, et al.
Published: (2025)
Qrazor: Reliable and Effortless 4-bit LLM Quantization by Significant Data Razoring
by: Lee, Dongyoung, et al.
Published: (2025)
by: Lee, Dongyoung, et al.
Published: (2025)
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
by: Park, Gunho, et al.
Published: (2025)
by: Park, Gunho, et al.
Published: (2025)
4bit-Quantization in Vector-Embedding for RAG
by: Jeong, Taehee
Published: (2025)
by: Jeong, Taehee
Published: (2025)
GalaxyDiT: Efficient Video Generation with Guidance Alignment and Adaptive Proxy in Diffusion Transformers
by: Song, Zhiye, et al.
Published: (2025)
by: Song, Zhiye, et al.
Published: (2025)
AAAC: Activation-Aware Adaptive Codebooks for 4-bit LLM Weight Quantization
by: IslamBouli, Beshr, et al.
Published: (2026)
by: IslamBouli, Beshr, et al.
Published: (2026)
TurboSAT: Gradient-Guided Boolean Satisfiability Accelerated on GPU-CPU Hybrid System
by: Dai, Steve, et al.
Published: (2025)
by: Dai, Steve, et al.
Published: (2025)
SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM Training
by: Jia, Jinda, et al.
Published: (2024)
by: Jia, Jinda, et al.
Published: (2024)
Revisiting Block-based Quantisation: What is Important for Sub-8-bit LLM Inference?
by: Zhang, Cheng, et al.
Published: (2023)
by: Zhang, Cheng, et al.
Published: (2023)
LoRaQ: Optimized Low Rank Approximation for 4-bit Quantization
by: Bouquet, Yann, et al.
Published: (2026)
by: Bouquet, Yann, et al.
Published: (2026)
ACE-RTL: When Agentic Context Evolution Meets RTL-Specialized LLMs
by: Deng, Chenhui, et al.
Published: (2026)
by: Deng, Chenhui, et al.
Published: (2026)
Fast and Efficient 2-bit LLM Inference on GPU: 2/4/16-bit in a Weight Matrix with Asynchronous Dequantization
by: Li, Jinhao, et al.
Published: (2023)
by: Li, Jinhao, et al.
Published: (2023)
Experts are all you need: A Composable Framework for Large Language Model Inference
by: Sridharan, Shrihari, et al.
Published: (2025)
by: Sridharan, Shrihari, et al.
Published: (2025)
GRPO with State Mutations: Improving LLM-Based Hardware Test Plan Generation
by: Kochar, Dimple Vijay, et al.
Published: (2026)
by: Kochar, Dimple Vijay, et al.
Published: (2026)
Atom: Low-bit Quantization for Efficient and Accurate LLM Serving
by: Zhao, Yilong, et al.
Published: (2023)
by: Zhao, Yilong, et al.
Published: (2023)
BitNet a4.8: 4-bit Activations for 1-bit LLMs
by: Wang, Hongyu, et al.
Published: (2024)
by: Wang, Hongyu, et al.
Published: (2024)
MergeQuant: Accurate 4-bit Static Quantization of Large Language Models by Channel-wise Calibration
by: Wang, Jinguang, et al.
Published: (2025)
by: Wang, Jinguang, et al.
Published: (2025)
DiRotQ: Rotation-Aware Quantization for 4-bit Diffusion Transformers
by: Sharify, Sayeh, et al.
Published: (2026)
by: Sharify, Sayeh, et al.
Published: (2026)
Block Rotation is All You Need for MXFP4 Quantization
by: Shao, Yuantian, et al.
Published: (2025)
by: Shao, Yuantian, et al.
Published: (2025)
Outlier Smoothing with Closed-Form Rotations for W4A4 Large Language Model Quantization
by: Xiao, Jinying, et al.
Published: (2025)
by: Xiao, Jinying, et al.
Published: (2025)
OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization
by: Li, Zhikai, et al.
Published: (2026)
by: Li, Zhikai, et al.
Published: (2026)
GradientSpace: Unsupervised Data Clustering for Improved Instruction Tuning
by: Sridharan, Shrihari, et al.
Published: (2025)
by: Sridharan, Shrihari, et al.
Published: (2025)
Pre-Registering the Detectable Effect: A Paired-MDE Budget for 4-bit Quantization Benchmarks, with a Pilot Audit
by: Zhuang, Zexin, et al.
Published: (2026)
by: Zhuang, Zexin, et al.
Published: (2026)
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
by: Lin, Yujun, et al.
Published: (2024)
by: Lin, Yujun, et al.
Published: (2024)
$R^2$-dLLM: Accelerating Diffusion Large Language Models via Spatio-Temporal Redundancy Reduction
by: Du, Zhenbang, et al.
Published: (2026)
by: Du, Zhenbang, et al.
Published: (2026)
Quantization-Aware Distillation for NVFP4 Inference Accuracy Recovery
by: Xin, Meng, et al.
Published: (2026)
by: Xin, Meng, et al.
Published: (2026)
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference
by: Jang, Wonsuk, et al.
Published: (2025)
by: Jang, Wonsuk, et al.
Published: (2025)
ICQuant: Index Coding enables Low-bit LLM Quantization
by: Li, Xinlin, et al.
Published: (2025)
by: Li, Xinlin, et al.
Published: (2025)
any4: Learned 4-bit Numeric Representation for LLMs
by: Elhoushi, Mostafa, et al.
Published: (2025)
by: Elhoushi, Mostafa, et al.
Published: (2025)
OSC: Hardware Efficient W4A4 Quantization via Outlier Separation in Channel Dimension
by: Zhang, Zhiyuan, et al.
Published: (2026)
by: Zhang, Zhiyuan, et al.
Published: (2026)
FireQ: Fast INT4-FP8 Kernel and RoPE-aware Quantization for LLM Inference Acceleration
by: Baek, Daehyeon, et al.
Published: (2025)
by: Baek, Daehyeon, et al.
Published: (2025)
SKIM: Any-bit Quantization Pushing The Limits of Post-Training Quantization
by: Bai, Runsheng, et al.
Published: (2024)
by: Bai, Runsheng, et al.
Published: (2024)
Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification
by: Pinckney, Nathaniel, et al.
Published: (2025)
by: Pinckney, Nathaniel, et al.
Published: (2025)
Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling
by: Cook, Jack, et al.
Published: (2025)
by: Cook, Jack, et al.
Published: (2025)
Similar Items
-
ESPACE: Dimensionality Reduction of Activations for Model Compression
by: Sakr, Charbel, et al.
Published: (2024) -
QuRL: Efficient Reinforcement Learning with Quantized Rollout
by: Li, Yuhang, et al.
Published: (2026) -
FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference
by: Hooper, Coleman, et al.
Published: (2025) -
ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning Models
by: Ramachandran, Akshat, et al.
Published: (2025) -
LLM4Cov: Execution-Aware Agentic Learning for High-coverage Testbench Generation
by: Zhang, Hejia, et al.
Published: (2026)