SURGE: Surrogate Gradient Adaptation in Binary Neural Networks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Huang, Haoyu, Liu, Boyu, Yang, Linlin, Li, Yanjing, Yang, Yuguang, Liu, Xuhui, Chen, Canyu, Fu, Zhongqian, Zhang, Baochang
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910252070338560
author Huang, Haoyu
Liu, Boyu
Yang, Linlin
Li, Yanjing
Yang, Yuguang
Liu, Xuhui
Chen, Canyu
Fu, Zhongqian
Zhang, Baochang
author_facet Huang, Haoyu
Liu, Boyu
Yang, Linlin
Li, Yanjing
Yang, Yuguang
Liu, Xuhui
Chen, Canyu
Fu, Zhongqian
Zhang, Baochang
contents The training of Binary Neural Networks (BNNs) is fundamentally based on gradient approximation for non-differentiable binarization operations (e.g., sign function). However, prevailing methods including the Straight-Through Estimator (STE) and its improved variants, rely on hand-crafted designs that suffer from gradient mismatch problem and information loss induced by fixed-range gradient clipping. To address this, we propose SURrogate GradiEnt Adaptation (SURGE), a novel learnable gradient compensation framework with theoretical grounding. SURGE mitigates gradient mismatch through auxiliary backpropagation. Specifically, we design a Dual-Path Gradient Compensator (DPGC) that constructs a parallel full-precision auxiliary branch for each binarized layer, decoupling gradient flow via output decomposition during backpropagation. DPGC enables bias-reduced gradient estimation by leveraging the full-precision branch to estimate components beyond STE's first-order approximation. To further enhance training stability, we introduce an Adaptive Gradient Scaler (AGS) based on an optimal scale factor to dynamically balance inter-branch gradient contributions via norm-based scaling. Experiments on image classification, object detection, and language understanding tasks demonstrate that SURGE performs best over state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10989
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SURGE: Surrogate Gradient Adaptation in Binary Neural Networks
Huang, Haoyu
Liu, Boyu
Yang, Linlin
Li, Yanjing
Yang, Yuguang
Liu, Xuhui
Chen, Canyu
Fu, Zhongqian
Zhang, Baochang
Machine Learning
Artificial Intelligence
The training of Binary Neural Networks (BNNs) is fundamentally based on gradient approximation for non-differentiable binarization operations (e.g., sign function). However, prevailing methods including the Straight-Through Estimator (STE) and its improved variants, rely on hand-crafted designs that suffer from gradient mismatch problem and information loss induced by fixed-range gradient clipping. To address this, we propose SURrogate GradiEnt Adaptation (SURGE), a novel learnable gradient compensation framework with theoretical grounding. SURGE mitigates gradient mismatch through auxiliary backpropagation. Specifically, we design a Dual-Path Gradient Compensator (DPGC) that constructs a parallel full-precision auxiliary branch for each binarized layer, decoupling gradient flow via output decomposition during backpropagation. DPGC enables bias-reduced gradient estimation by leveraging the full-precision branch to estimate components beyond STE's first-order approximation. To further enhance training stability, we introduce an Adaptive Gradient Scaler (AGS) based on an optimal scale factor to dynamically balance inter-branch gradient contributions via norm-based scaling. Experiments on image classification, object detection, and language understanding tasks demonstrate that SURGE performs best over state-of-the-art methods.
title SURGE: Surrogate Gradient Adaptation in Binary Neural Networks
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.10989