DeInfoReg: A Decoupled Learning Framework for Better Training Throughput

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Zih-Hao, Lin, You-Teng, Chen, Hung-Hsuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913942266183680
author Huang, Zih-Hao
Lin, You-Teng
Chen, Hung-Hsuan
author_facet Huang, Zih-Hao
Lin, You-Teng
Chen, Hung-Hsuan
contents This paper introduces Decoupled Supervised Learning with Information Regularization (DeInfoReg), a novel approach that transforms a long gradient flow into multiple shorter ones, thereby mitigating the vanishing gradient problem. Integrating a pipeline strategy, DeInfoReg enables model parallelization across multiple GPUs, significantly improving training throughput. We compare our proposed method with standard backpropagation and other gradient flow decomposition techniques. Extensive experiments on diverse tasks and datasets demonstrate that DeInfoReg achieves superior performance and better noise resistance than traditional BP models and efficiently utilizes parallel computing resources. The code for reproducibility is available at: https://github.com/ianzih/Decoupled-Supervised-Learning-for-Information-Regularization/.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18193
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DeInfoReg: A Decoupled Learning Framework for Better Training Throughput
Huang, Zih-Hao
Lin, You-Teng
Chen, Hung-Hsuan
Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
This paper introduces Decoupled Supervised Learning with Information Regularization (DeInfoReg), a novel approach that transforms a long gradient flow into multiple shorter ones, thereby mitigating the vanishing gradient problem. Integrating a pipeline strategy, DeInfoReg enables model parallelization across multiple GPUs, significantly improving training throughput. We compare our proposed method with standard backpropagation and other gradient flow decomposition techniques. Extensive experiments on diverse tasks and datasets demonstrate that DeInfoReg achieves superior performance and better noise resistance than traditional BP models and efficiently utilizes parallel computing resources. The code for reproducibility is available at: https://github.com/ianzih/Decoupled-Supervised-Learning-for-Information-Regularization/.
title DeInfoReg: A Decoupled Learning Framework for Better Training Throughput
topic Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2506.18193