BackSlash: Rate Constrained Optimized Training of Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wu, Jun, Wen, Jiangtao, Han, Yuxing
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918207654199296
author Wu, Jun
Wen, Jiangtao
Han, Yuxing
author_facet Wu, Jun
Wen, Jiangtao
Han, Yuxing
contents The rapid advancement of large-language models (LLMs) has driven extensive research into parameter compression after training has been completed, yet compression during the training phase remains largely unexplored. In this work, we introduce Rate-Constrained Training (BackSlash), a novel training-time compression approach based on rate-distortion optimization (RDO). BackSlash enables a flexible trade-off between model accuracy and complexity, significantly reducing parameter redundancy while preserving performance. Experiments in various architectures and tasks demonstrate that BackSlash can reduce memory usage by 60% - 90% without accuracy loss and provides significant compression gain compared to compression after training. Moreover, BackSlash proves to be highly versatile: it enhances generalization with small Lagrange multipliers, improves model robustness to pruning (maintaining accuracy even at 80% pruning rates), and enables network simplification for accelerated inference on edge devices.
format Preprint
id arxiv_https___arxiv_org_abs_2504_16968
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BackSlash: Rate Constrained Optimized Training of Large Language Models
Wu, Jun
Wen, Jiangtao
Han, Yuxing
Machine Learning
Artificial Intelligence
The rapid advancement of large-language models (LLMs) has driven extensive research into parameter compression after training has been completed, yet compression during the training phase remains largely unexplored. In this work, we introduce Rate-Constrained Training (BackSlash), a novel training-time compression approach based on rate-distortion optimization (RDO). BackSlash enables a flexible trade-off between model accuracy and complexity, significantly reducing parameter redundancy while preserving performance. Experiments in various architectures and tasks demonstrate that BackSlash can reduce memory usage by 60% - 90% without accuracy loss and provides significant compression gain compared to compression after training. Moreover, BackSlash proves to be highly versatile: it enhances generalization with small Lagrange multipliers, improves model robustness to pruning (maintaining accuracy even at 80% pruning rates), and enables network simplification for accelerated inference on edge devices.
title BackSlash: Rate Constrained Optimized Training of Large Language Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2504.16968