Learning Discretized Neural Networks under Ricci Flow

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Jun, Chen, Hanwen, Wang, Mengmeng, Dai, Guang, Tsang, Ivor W., Liu, Yong
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916531286310912
author Chen, Jun
Chen, Hanwen
Wang, Mengmeng
Dai, Guang
Tsang, Ivor W.
Liu, Yong
author_facet Chen, Jun
Chen, Hanwen
Wang, Mengmeng
Dai, Guang
Tsang, Ivor W.
Liu, Yong
contents In this paper, we study Discretized Neural Networks (DNNs) composed of low-precision weights and activations, which suffer from either infinite or zero gradients due to the non-differentiable discrete function during training. Most training-based DNNs in such scenarios employ the standard Straight-Through Estimator (STE) to approximate the gradient w.r.t. discrete values. However, the use of STE introduces the problem of gradient mismatch, arising from perturbations in the approximated gradient. To address this problem, this paper reveals that this mismatch can be interpreted as a metric perturbation in a Riemannian manifold, viewed through the lens of duality theory. Building on information geometry, we construct the Linearly Nearly Euclidean (LNE) manifold for DNNs, providing a background for addressing perturbations. By introducing a partial differential equation on metrics, i.e., the Ricci flow, we establish the dynamical stability and convergence of the LNE metric with the $L^2$-norm perturbation. In contrast to previous perturbation theories with convergence rates in fractional powers, the metric perturbation under the Ricci flow exhibits exponential decay in the LNE manifold. Experimental results across various datasets demonstrate that our method achieves superior and more stable performance for DNNs compared to other representative training-based methods.
format Preprint
id arxiv_https___arxiv_org_abs_2302_03390
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Learning Discretized Neural Networks under Ricci Flow
Chen, Jun
Chen, Hanwen
Wang, Mengmeng
Dai, Guang
Tsang, Ivor W.
Liu, Yong
Machine Learning
Information Theory
Neural and Evolutionary Computing
In this paper, we study Discretized Neural Networks (DNNs) composed of low-precision weights and activations, which suffer from either infinite or zero gradients due to the non-differentiable discrete function during training. Most training-based DNNs in such scenarios employ the standard Straight-Through Estimator (STE) to approximate the gradient w.r.t. discrete values. However, the use of STE introduces the problem of gradient mismatch, arising from perturbations in the approximated gradient. To address this problem, this paper reveals that this mismatch can be interpreted as a metric perturbation in a Riemannian manifold, viewed through the lens of duality theory. Building on information geometry, we construct the Linearly Nearly Euclidean (LNE) manifold for DNNs, providing a background for addressing perturbations. By introducing a partial differential equation on metrics, i.e., the Ricci flow, we establish the dynamical stability and convergence of the LNE metric with the $L^2$-norm perturbation. In contrast to previous perturbation theories with convergence rates in fractional powers, the metric perturbation under the Ricci flow exhibits exponential decay in the LNE manifold. Experimental results across various datasets demonstrate that our method achieves superior and more stable performance for DNNs compared to other representative training-based methods.
title Learning Discretized Neural Networks under Ricci Flow
topic Machine Learning
Information Theory
Neural and Evolutionary Computing
url https://arxiv.org/abs/2302.03390