LoopQ: Quantization for Recursive Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Fang, Rui, Chen, Hsi-Wen, Chen, Ming-Syan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Amortized-Precision Quantization for Early-Exit Vision Transformers
by: Fang, Rui, et al.
Published: (2026)
by: Fang, Rui, et al.
Published: (2026)
KV Admission: Learning What to Write for Efficient Long-Context Inference
by: Huang, Yen-Chieh, et al.
Published: (2025)
by: Huang, Yen-Chieh, et al.
Published: (2025)
CiMRAG: CiM-Aware Domain-Adaptive and Noise-Resilient Retrieval-Augmented Generation for Edge-Based LLMs
by: Chiu, Shih-Hsuan, et al.
Published: (2026)
by: Chiu, Shih-Hsuan, et al.
Published: (2026)
Local Shapley: Model-Induced Locality and Optimal Reuse in Data Valuation
by: Yang, Xuan, et al.
Published: (2026)
by: Yang, Xuan, et al.
Published: (2026)
Looping Back to Move Forward: Recursive Transformers for Efficient and Flexible Large Multimodal Models
by: Xu, Ruihan, et al.
Published: (2026)
by: Xu, Ruihan, et al.
Published: (2026)
Interference-Aware Multi-Task Unlearning
by: Huang, Ying-Hua, et al.
Published: (2026)
by: Huang, Ying-Hua, et al.
Published: (2026)
What Makes Looped Transformers Perform Better Than Non-Recursive Ones
by: Gong, Zixuan, et al.
Published: (2025)
by: Gong, Zixuan, et al.
Published: (2025)
Recursive Backwards Q-Learning in Deterministic Environments
by: Diekhoff, Jan, et al.
Published: (2024)
by: Diekhoff, Jan, et al.
Published: (2024)
Q-resafe: Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language Models
by: Chen, Kejia, et al.
Published: (2025)
by: Chen, Kejia, et al.
Published: (2025)
MeSH: Memory-as-State-Highways for Recursive Transformers
by: Yu, Chengting, et al.
Published: (2025)
by: Yu, Chengting, et al.
Published: (2025)
Unlocking Out-of-Distribution Generalization in Transformers via Recursive Latent Space Reasoning
by: Altabaa, Awni, et al.
Published: (2025)
by: Altabaa, Awni, et al.
Published: (2025)
CoreQ: Learning-Free Mismatch Correction and Successive Rounding for Quantization
by: Cha, Seohyeon, et al.
Published: (2026)
by: Cha, Seohyeon, et al.
Published: (2026)
Pioneering 4-Bit FP Quantization for Diffusion Models: Mixup-Sign Quantization and Timestep-Aware Fine-Tuning
by: Zhao, Maosen, et al.
Published: (2025)
by: Zhao, Maosen, et al.
Published: (2025)
Q-MMR: Off-Policy Evaluation via Recursive Reweighting and Moment Matching
by: Li, Xiang, et al.
Published: (2026)
by: Li, Xiang, et al.
Published: (2026)
FIMA-Q: Post-Training Quantization for Vision Transformers by Fisher Information Matrix Approximation
by: Wu, Zhuguanyu, et al.
Published: (2025)
by: Wu, Zhuguanyu, et al.
Published: (2025)
Latent Flow Transformer
by: Wu, Yen-Chen, et al.
Published: (2025)
by: Wu, Yen-Chen, et al.
Published: (2025)
On the Design Space Between Transformers and Recursive Neural Nets
by: Chowdhury, Jishnu Ray, et al.
Published: (2024)
by: Chowdhury, Jishnu Ray, et al.
Published: (2024)
Stability and Generalization in Looped Transformers
by: Labovich, Asher
Published: (2026)
by: Labovich, Asher
Published: (2026)
GlowQ: Group-Shared LOw-Rank Approximation for Quantized LLMs
by: An, Selim, et al.
Published: (2026)
by: An, Selim, et al.
Published: (2026)
Q-realign: Piggybacking Realignment on Quantization for Safe and Efficient LLM Deployment
by: Tan, Qitao, et al.
Published: (2026)
by: Tan, Qitao, et al.
Published: (2026)
To CoT or To Loop? A Formal Comparison Between Chain-of-Thought and Looped Transformers
by: Xu, Kevin, et al.
Published: (2025)
by: Xu, Kevin, et al.
Published: (2025)
Calibration and Transformation-Free Weight-Only LLMs Quantization via Dynamic Grouping
by: Zheng, Xinzhe, et al.
Published: (2025)
by: Zheng, Xinzhe, et al.
Published: (2025)
Bypassing the Exponential Dependency: Looped Transformers Efficiently Learn In-context by Multi-step Gradient Descent
by: Chen, Bo, et al.
Published: (2024)
by: Chen, Bo, et al.
Published: (2024)
Navigation with QPHIL: Quantizing Planner for Hierarchical Implicit Q-Learning
by: Canesse, Alexi, et al.
Published: (2024)
by: Canesse, Alexi, et al.
Published: (2024)
Sparse-VQ Transformer: An FFN-Free Framework with Vector Quantization for Enhanced Time Series Forecasting
by: Zhao, Yanjun, et al.
Published: (2024)
by: Zhao, Yanjun, et al.
Published: (2024)
HeadQ: Model-Visible Distortion and Score-Space Correction for KV-Cache Quantization
by: Williams, Jorge L. Ruiz
Published: (2026)
by: Williams, Jorge L. Ruiz
Published: (2026)
RWKVQuant: Quantizing the RWKV Family with Proxy Guided Hybrid of Scalar and Vector Quantization
by: Xu, Chen, et al.
Published: (2025)
by: Xu, Chen, et al.
Published: (2025)
On-the-Fly Adaptation to Quantization: Configuration-Aware LoRA for Efficient Fine-Tuning of Quantized LLMs
by: Ye, Rongguang, et al.
Published: (2025)
by: Ye, Rongguang, et al.
Published: (2025)
OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM Reasoning
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
Relational Preference Encoding in Looped Transformer Internal States
by: Kirin, Jan
Published: (2026)
by: Kirin, Jan
Published: (2026)
PolarQuant: Quantizing KV Caches with Polar Transformation
by: Han, Insu, et al.
Published: (2025)
by: Han, Insu, et al.
Published: (2025)
AutoMixQ: Self-Adjusting Quantization for High Performance Memory-Efficient Fine-Tuning
by: Zhou, Changhai, et al.
Published: (2024)
by: Zhou, Changhai, et al.
Published: (2024)
CLoQ: Enhancing Fine-Tuning of Quantized LLMs via Calibrated LoRA Initialization
by: Deng, Yanxia, et al.
Published: (2025)
by: Deng, Yanxia, et al.
Published: (2025)
Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM Deployment
by: Lee, Deokjae, et al.
Published: (2025)
by: Lee, Deokjae, et al.
Published: (2025)
OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting
by: Hu, Xing, et al.
Published: (2025)
by: Hu, Xing, et al.
Published: (2025)
QFlash: Bridging Quantization and Memory Efficiency in Vision Transformer Attention
by: Oh, Sehyeon, et al.
Published: (2026)
by: Oh, Sehyeon, et al.
Published: (2026)
On-device AI: Quantization-aware Training of Transformers in Time-Series
by: Ling, Tianheng, et al.
Published: (2024)
by: Ling, Tianheng, et al.
Published: (2024)
HeRo-Q: A General Framework for Stable Low Bit Quantization via Hessian Conditioning
by: Zhang, Jinhao Zhang Yunquan, et al.
Published: (2026)
by: Zhang, Jinhao Zhang Yunquan, et al.
Published: (2026)
Q-ROAR: Outlier-Aware Rescaling for RoPE Position Interpolation in Quantized Long-Context LLMs
by: Qiao, Ye, et al.
Published: (2025)
by: Qiao, Ye, et al.
Published: (2025)
Quantizing Text-attributed Graphs for Semantic-Structural Integration
by: Bo, Jianyuan, et al.
Published: (2025)
by: Bo, Jianyuan, et al.
Published: (2025)
Similar Items
-
Amortized-Precision Quantization for Early-Exit Vision Transformers
by: Fang, Rui, et al.
Published: (2026) -
KV Admission: Learning What to Write for Efficient Long-Context Inference
by: Huang, Yen-Chieh, et al.
Published: (2025) -
CiMRAG: CiM-Aware Domain-Adaptive and Noise-Resilient Retrieval-Augmented Generation for Edge-Based LLMs
by: Chiu, Shih-Hsuan, et al.
Published: (2026) -
Local Shapley: Model-Induced Locality and Optimal Reuse in Data Valuation
by: Yang, Xuan, et al.
Published: (2026) -
Looping Back to Move Forward: Recursive Transformers for Efficient and Flexible Large Multimodal Models
by: Xu, Ruihan, et al.
Published: (2026)