On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gupta, Aarav, Deshpande, Gururaj, Chakraborty, Chandreyi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910156866977792
author Gupta, Aarav
Deshpande, Gururaj
Chakraborty, Chandreyi
author_facet Gupta, Aarav
Deshpande, Gururaj
Chakraborty, Chandreyi
contents Auto-regressive Large Language Models (LLMs) achieve strong performance on coding tasks, but incur high memory and inference costs. Diffusion-based language models (d-LLMs) offer bounded inference cost via iterative denoising, but their behavior under post-training quantization (PTQ) has been sparsely explored. We investigate the application and robustness of PTQ techniques, specifically GPTQ and a modified Hessian-Aware Quantization (HAWQ) algorithm, on a diffusion-based coding LLM (CoDA) and observe that these methods applied to CoDA exhibit greater robustness at low bitwidths compared to Qwen3-1.7B, its auto-regressive counterpart, under a standardized evaluation pipeline. We find that in our setup, CoDA exhibits greater robustness at low bitwidths (2-4 bits), with smaller accuracy degradation across HumanEval and MBPP benchmarks. Additionally, mixed-precision configurations derived from HAWQ provide smooth trade-offs across accuracy, latency, and memory. The results suggest that diffusion LLMs may offer advantages for efficient deployment due to more quantization-resilience.
format Preprint
id arxiv_https___arxiv_org_abs_2604_20079
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks
Gupta, Aarav
Deshpande, Gururaj
Chakraborty, Chandreyi
Machine Learning
Computation and Language
Auto-regressive Large Language Models (LLMs) achieve strong performance on coding tasks, but incur high memory and inference costs. Diffusion-based language models (d-LLMs) offer bounded inference cost via iterative denoising, but their behavior under post-training quantization (PTQ) has been sparsely explored. We investigate the application and robustness of PTQ techniques, specifically GPTQ and a modified Hessian-Aware Quantization (HAWQ) algorithm, on a diffusion-based coding LLM (CoDA) and observe that these methods applied to CoDA exhibit greater robustness at low bitwidths compared to Qwen3-1.7B, its auto-regressive counterpart, under a standardized evaluation pipeline. We find that in our setup, CoDA exhibits greater robustness at low bitwidths (2-4 bits), with smaller accuracy degradation across HumanEval and MBPP benchmarks. Additionally, mixed-precision configurations derived from HAWQ provide smooth trade-offs across accuracy, latency, and memory. The results suggest that diffusion LLMs may offer advantages for efficient deployment due to more quantization-resilience.
title On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2604.20079