LoopQ: Quantization for Recursive Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fang, Rui, Chen, Hsi-Wen, Chen, Ming-Syan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909048508514304
author Fang, Rui
Chen, Hsi-Wen
Chen, Ming-Syan
author_facet Fang, Rui
Chen, Hsi-Wen
Chen, Ming-Syan
contents Looped language models (LoopLMs) improve parameter efficiency by recursively reusing Transformer blocks, enabling deeper computation under a fixed model size. However, this reuse makes LoopLMs more fragile under post-training quantization (PTQ). We present the first systematic study of quantization in LoopLMs and identify three challenges: distribution shift across roles, state reuse across loop transitions, and recursive error accumulation. To address these challenges, we propose LoopQ, a loop-aware PTQ framework that preserves a shared quantized backbone while introducing lightweight adaptations. LoopQ combines activation scaling, selective transformation, cross-loop state alignment, and trajectory-aware optimization to reduce distributional mismatch within loops and error accumulation across loops. Experiments across seven benchmarks show that, under W4A4 quantization, LoopQ improves average downstream accuracy by 68.8% and reduces average perplexity by 87.7% compared with the strongest static PTQ baseline.
format Preprint
id arxiv_https___arxiv_org_abs_2605_16343
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LoopQ: Quantization for Recursive Transformers
Fang, Rui
Chen, Hsi-Wen
Chen, Ming-Syan
Machine Learning
Artificial Intelligence
Looped language models (LoopLMs) improve parameter efficiency by recursively reusing Transformer blocks, enabling deeper computation under a fixed model size. However, this reuse makes LoopLMs more fragile under post-training quantization (PTQ). We present the first systematic study of quantization in LoopLMs and identify three challenges: distribution shift across roles, state reuse across loop transitions, and recursive error accumulation. To address these challenges, we propose LoopQ, a loop-aware PTQ framework that preserves a shared quantized backbone while introducing lightweight adaptations. LoopQ combines activation scaling, selective transformation, cross-loop state alignment, and trajectory-aware optimization to reduce distributional mismatch within loops and error accumulation across loops. Experiments across seven benchmarks show that, under W4A4 quantization, LoopQ improves average downstream accuracy by 68.8% and reduces average perplexity by 87.7% compared with the strongest static PTQ baseline.
title LoopQ: Quantization for Recursive Transformers
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.16343