FlexQ: Efficient Post-training INT6 Quantization for LLM Serving via Algorithm-System Co-Design
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Hao, Jia, Aining, Bu, Weifeng, Cai, Yushu, Sheng, Kai, Chen, Hao, He, Xin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SAW-INT4: System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving
by: Jia, Jinda, et al.
Published: (2026)
by: Jia, Jinda, et al.
Published: (2026)
FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
by: Oliaro, Gabriele, et al.
Published: (2024)
by: Oliaro, Gabriele, et al.
Published: (2024)
Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference
by: Wang, Xinyu, et al.
Published: (2026)
by: Wang, Xinyu, et al.
Published: (2026)
Efficient Post-training Quantization with FP8 Formats
by: Shen, Haihao, et al.
Published: (2023)
by: Shen, Haihao, et al.
Published: (2023)
INT-FlashAttention: Enabling Flash Attention for INT8 Quantization
by: Chen, Shimao, et al.
Published: (2024)
by: Chen, Shimao, et al.
Published: (2024)
FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design
by: Xia, Haojun, et al.
Published: (2024)
by: Xia, Haojun, et al.
Published: (2024)
FireQ: Fast INT4-FP8 Kernel and RoPE-aware Quantization for LLM Inference Acceleration
by: Baek, Daehyeon, et al.
Published: (2025)
by: Baek, Daehyeon, et al.
Published: (2025)
Jetfire: Efficient and Accurate Transformer Pretraining with INT8 Data Flow and Per-Block Quantization
by: Xi, Haocheng, et al.
Published: (2024)
by: Xi, Haocheng, et al.
Published: (2024)
SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization
by: Zhang, Jintao, et al.
Published: (2024)
by: Zhang, Jintao, et al.
Published: (2024)
Q-VLM: Post-training Quantization for Large Vision-Language Models
by: Wang, Changyuan, et al.
Published: (2024)
by: Wang, Changyuan, et al.
Published: (2024)
FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization
by: Liu, Fangxin, et al.
Published: (2025)
by: Liu, Fangxin, et al.
Published: (2025)
PersonalQ: Select, Quantize, and Serve Personalized Diffusion Models for Efficient Inference
by: Wang, Qirui, et al.
Published: (2026)
by: Wang, Qirui, et al.
Published: (2026)
Q-GaLore: Quantized GaLore with INT4 Projection and Layer-Adaptive Low-Rank Gradients
by: Zhang, Zhenyu, et al.
Published: (2024)
by: Zhang, Zhenyu, et al.
Published: (2024)
FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation
by: Wu, Yinpeng, et al.
Published: (2026)
by: Wu, Yinpeng, et al.
Published: (2026)
Atom: Low-bit Quantization for Efficient and Accurate LLM Serving
by: Zhao, Yilong, et al.
Published: (2023)
by: Zhao, Yilong, et al.
Published: (2023)
AgentServe: Algorithm-System Co-Design for Efficient Agentic AI Serving on a Consumer-Grade GPU
by: Zhang, Yuning, et al.
Published: (2026)
by: Zhang, Yuning, et al.
Published: (2026)
QVD: Post-training Quantization for Video Diffusion Models
by: Tian, Shilong, et al.
Published: (2024)
by: Tian, Shilong, et al.
Published: (2024)
MoEless: Efficient MoE LLM Serving via Serverless Computing
by: Yu, Hanfei, et al.
Published: (2026)
by: Yu, Hanfei, et al.
Published: (2026)
PQD: Post-training Quantization for Efficient Diffusion Models
by: Ye, Jiaojiao, et al.
Published: (2024)
by: Ye, Jiaojiao, et al.
Published: (2024)
Progressive Sparse Attention: Algorithm and System Co-design for Efficient Attention in LLM Serving
by: Zhou, Qihui, et al.
Published: (2025)
by: Zhou, Qihui, et al.
Published: (2025)
CARD: A Cache-Assisted Parallel Speculative Decoding Framework via Query-and-Correct Paradigm for Accelerating LLM Inference
by: Zhou, Enyu, et al.
Published: (2025)
by: Zhou, Enyu, et al.
Published: (2025)
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
by: Lin, Yujun, et al.
Published: (2024)
by: Lin, Yujun, et al.
Published: (2024)
FlexLLM: Composable HLS Library for Flexible Hybrid LLM Accelerator Design
by: Zhang, Jiahao, et al.
Published: (2026)
by: Zhang, Jiahao, et al.
Published: (2026)
FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization
by: Lee, Jung Hyun, et al.
Published: (2023)
by: Lee, Jung Hyun, et al.
Published: (2023)
FF-INT8: Efficient Forward-Forward DNN Training on Edge Devices with INT8 Precision
by: Ma, Jingxiao, et al.
Published: (2025)
by: Ma, Jingxiao, et al.
Published: (2025)
Quantum Amplitude‐Phase Judgment Circuits for Full Quantization Algorithms
by: Ziming Dong, et al.
Published: (2025)
by: Ziming Dong, et al.
Published: (2025)
FlexQuant: Elastic Quantization Framework for Locally Hosted LLM on Edge Devices
by: Chai, Yuji, et al.
Published: (2025)
by: Chai, Yuji, et al.
Published: (2025)
QwT-v2: Practical, Effective and Efficient Post-Training Quantization
by: Tang, Ningyuan, et al.
Published: (2025)
by: Tang, Ningyuan, et al.
Published: (2025)
Efficient INT8 Single-Image Super-Resolution via Deployment-Aware Quantization and Teacher-Guided Training
by: Nguyen, Pham Phuong Nam, et al.
Published: (2026)
by: Nguyen, Pham Phuong Nam, et al.
Published: (2026)
Q-DiT: Accurate Post-Training Quantization for Diffusion Transformers
by: Chen, Lei, et al.
Published: (2024)
by: Chen, Lei, et al.
Published: (2024)
Towards Accurate Post-training Quantization for Reparameterized Models
by: Zhang, Luoming, et al.
Published: (2024)
by: Zhang, Luoming, et al.
Published: (2024)
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
by: Su, Zhaoyuan, et al.
Published: (2025)
by: Su, Zhaoyuan, et al.
Published: (2025)
Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization
by: Kim, Minsu, et al.
Published: (2025)
by: Kim, Minsu, et al.
Published: (2025)
COMQ: A Backpropagation-Free Algorithm for Post-Training Quantization
by: Zhang, Aozhong, et al.
Published: (2024)
by: Zhang, Aozhong, et al.
Published: (2024)
FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
by: Lin, Yanying, et al.
Published: (2025)
by: Lin, Yanying, et al.
Published: (2025)
TesseraQ: Ultra Low-Bit LLM Post-Training Quantization with Block Reconstruction
by: Li, Yuhang, et al.
Published: (2024)
by: Li, Yuhang, et al.
Published: (2024)
Multi-stage Flow Scheduling for LLM Serving
by: Sun, Yijun, et al.
Published: (2026)
by: Sun, Yijun, et al.
Published: (2026)
Q&C: When Quantization Meets Cache in Efficient Image Generation
by: Ding, Xin, et al.
Published: (2025)
by: Ding, Xin, et al.
Published: (2025)
GPU-Accelerated INT8 Quantization for KV Cache Compression in Large Language Models
by: Taneja, Maanas, et al.
Published: (2026)
by: Taneja, Maanas, et al.
Published: (2026)
Toward INT4 Fixed-Point Training via Exploring Quantization Error for Gradients
by: Kim, Dohyung, et al.
Published: (2024)
by: Kim, Dohyung, et al.
Published: (2024)
Similar Items
-
SAW-INT4: System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving
by: Jia, Jinda, et al.
Published: (2026) -
FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
by: Oliaro, Gabriele, et al.
Published: (2024) -
Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference
by: Wang, Xinyu, et al.
Published: (2026) -
Efficient Post-training Quantization with FP8 Formats
by: Shen, Haihao, et al.
Published: (2023) -
INT-FlashAttention: Enabling Flash Attention for INT8 Quantization
by: Chen, Shimao, et al.
Published: (2024)