Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Lu, Haiquan, Chen, Zigeng, Fang, Gongfan, Ma, Xinyin, Wang, Xinchao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
dParallel: Learnable Parallel Decoding for dLLMs
by: Chen, Zigeng, et al.
Published: (2025)
by: Chen, Zigeng, et al.
Published: (2025)
dMoE: dLLMs with Learnable Block Experts
by: Feng, Sicheng, et al.
Published: (2026)
by: Feng, Sicheng, et al.
Published: (2026)
dVoting: Fast Voting for dLLMs
by: Feng, Sicheng, et al.
Published: (2026)
by: Feng, Sicheng, et al.
Published: (2026)
MixReasoning: Switching Modes to Think
by: Lu, Haiquan, et al.
Published: (2025)
by: Lu, Haiquan, et al.
Published: (2025)
Collaborative Decoding Makes Visual Auto-Regressive Modeling Efficient
by: Chen, Zigeng, et al.
Published: (2024)
by: Chen, Zigeng, et al.
Published: (2024)
Thinkless: LLM Learns When to Think
by: Fang, Gongfan, et al.
Published: (2025)
by: Fang, Gongfan, et al.
Published: (2025)
DMax: Aggressive Parallel Decoding for dLLMs
by: Chen, Zigeng, et al.
Published: (2026)
by: Chen, Zigeng, et al.
Published: (2026)
SlimSAM: 0.1% Data Makes Segment Anything Slim
by: Chen, Zigeng, et al.
Published: (2023)
by: Chen, Zigeng, et al.
Published: (2023)
dKV-Cache: The Cache for Diffusion Language Models
by: Ma, Xinyin, et al.
Published: (2025)
by: Ma, Xinyin, et al.
Published: (2025)
Efficient Reasoning Models: A Survey
by: Feng, Sicheng, et al.
Published: (2025)
by: Feng, Sicheng, et al.
Published: (2025)
Remix-DiT: Mixing Diffusion Transformers for Multi-Expert Denoising
by: Fang, Gongfan, et al.
Published: (2024)
by: Fang, Gongfan, et al.
Published: (2024)
ConciseHint: Boosting Efficient Reasoning via Continuous Concise Hints during Generation
by: Tang, Siao, et al.
Published: (2025)
by: Tang, Siao, et al.
Published: (2025)
AsyncDiff: Parallelizing Diffusion Models by Asynchronous Denoising
by: Chen, Zigeng, et al.
Published: (2024)
by: Chen, Zigeng, et al.
Published: (2024)
SparseD: Sparse Attention for Diffusion Language Models
by: Wang, Zeqing, et al.
Published: (2025)
by: Wang, Zeqing, et al.
Published: (2025)
CoT-Valve: Length-Compressible Chain-of-Thought Tuning
by: Ma, Xinyin, et al.
Published: (2025)
by: Ma, Xinyin, et al.
Published: (2025)
In-Video Instructions: Visual Signals as Generative Control
by: Fang, Gongfan, et al.
Published: (2025)
by: Fang, Gongfan, et al.
Published: (2025)
VeriThinker: Learning to Verify Makes Reasoning Model Efficient
by: Chen, Zigeng, et al.
Published: (2025)
by: Chen, Zigeng, et al.
Published: (2025)
Q-ARVD: Quantizing Autoregressive Video Diffusion Models
by: Tang, Siao, et al.
Published: (2026)
by: Tang, Siao, et al.
Published: (2026)
Every Step Counts: Decoding Trajectories as Authorship Fingerprints of dLLMs
by: Li, Qi, et al.
Published: (2025)
by: Li, Qi, et al.
Published: (2025)
TinyFusion: Diffusion Transformers Learned Shallow
by: Fang, Gongfan, et al.
Published: (2024)
by: Fang, Gongfan, et al.
Published: (2024)
Learning-to-Cache: Accelerating Diffusion Transformer via Layer Caching
by: Ma, Xinyin, et al.
Published: (2024)
by: Ma, Xinyin, et al.
Published: (2024)
Isomorphic Pruning for Vision Models
by: Fang, Gongfan, et al.
Published: (2024)
by: Fang, Gongfan, et al.
Published: (2024)
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization
by: Tao, Wei, et al.
Published: (2026)
by: Tao, Wei, et al.
Published: (2026)
Invisible Safety Threat: Malicious Finetuning for LLM via Steganography
by: Wan, Guangnian, et al.
Published: (2026)
by: Wan, Guangnian, et al.
Published: (2026)
LiteFocus: Accelerated Diffusion Inference for Long Audio Synthesis
by: Tan, Zhenxiong, et al.
Published: (2024)
by: Tan, Zhenxiong, et al.
Published: (2024)
Dimple: Discrete Diffusion Multimodal Large Language Model with Parallel Decoding
by: Yu, Runpeng, et al.
Published: (2025)
by: Yu, Runpeng, et al.
Published: (2025)
RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory
by: Zuo, Fei, et al.
Published: (2026)
by: Zuo, Fei, et al.
Published: (2026)
Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge
by: Shen, Xuan, et al.
Published: (2023)
by: Shen, Xuan, et al.
Published: (2023)
GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs
by: Deng, Jianing, et al.
Published: (2026)
by: Deng, Jianing, et al.
Published: (2026)
Self-Purification Mitigates Backdoors in Multimodal Diffusion Language Models
by: Wan, Guangnian, et al.
Published: (2026)
by: Wan, Guangnian, et al.
Published: (2026)
MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM
by: Wang, Dongwei, et al.
Published: (2026)
by: Wang, Dongwei, et al.
Published: (2026)
Accelerating Prefilling via Decoding-time Contribution Sparsity
by: He, Zhiyuan, et al.
Published: (2025)
by: He, Zhiyuan, et al.
Published: (2025)
PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration
by: Wu, Songhao, et al.
Published: (2025)
by: Wu, Songhao, et al.
Published: (2025)
DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs
by: Lin, Haokun, et al.
Published: (2024)
by: Lin, Haokun, et al.
Published: (2024)
PDTrim: Targeted Pruning for Prefill-Decode Disaggregation in Inference
by: Zhang, Hao, et al.
Published: (2025)
by: Zhang, Hao, et al.
Published: (2025)
FlatQuant: Flatness Matters for LLM Quantization
by: Sun, Yuxuan, et al.
Published: (2024)
by: Sun, Yuxuan, et al.
Published: (2024)
ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
by: Liang, Yesheng, et al.
Published: (2025)
by: Liang, Yesheng, et al.
Published: (2025)
DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization
by: Shao, Yuantian, et al.
Published: (2025)
by: Shao, Yuantian, et al.
Published: (2025)
MobileQuant: Mobile-friendly Quantization for On-device Language Models
by: Tan, Fuwen, et al.
Published: (2024)
by: Tan, Fuwen, et al.
Published: (2024)
Copy-as-Decode: Grammar-Constrained Parallel Prefill for LLM Editing
by: Liu, Ziyang
Published: (2026)
by: Liu, Ziyang
Published: (2026)
Similar Items
-
dParallel: Learnable Parallel Decoding for dLLMs
by: Chen, Zigeng, et al.
Published: (2025) -
dMoE: dLLMs with Learnable Block Experts
by: Feng, Sicheng, et al.
Published: (2026) -
dVoting: Fast Voting for dLLMs
by: Feng, Sicheng, et al.
Published: (2026) -
MixReasoning: Switching Modes to Think
by: Lu, Haiquan, et al.
Published: (2025) -
Collaborative Decoding Makes Visual Auto-Regressive Modeling Efficient
by: Chen, Zigeng, et al.
Published: (2024)