Improving Discrete Optimisation Via Decoupled Straight-Through Estimator

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Shah, Rushi, Yan, Mingyuan, Mozer, Michael Curtis, Liu, Dianbo
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910028355600384
author Shah, Rushi
Yan, Mingyuan
Mozer, Michael Curtis
Liu, Dianbo
author_facet Shah, Rushi
Yan, Mingyuan
Mozer, Michael Curtis
Liu, Dianbo
contents The Straight-Through Estimator (STE) is the dominant method for training neural networks with discrete variables, enabling gradient-based optimisation by routing gradients through a differentiable surrogate. However, existing STE variants conflate two fundamentally distinct concerns: forward-pass stochasticity, which controls exploration and latent space utilisation, and backward-pass gradient dispersion i.e how learning signals are distributed across categories. We show that these concerns are qualitatively different and that tying them to a single temperature parameter leaves significant performance gains untapped. We propose Decoupled Straight-Through (Decoupled ST), a minimal modification that introduces separate temperatures for the forward pass ($τ_f$) and the backward pass ($τ_b$). This simple change enables independent tuning of exploration and gradient dispersion. Across three diverse tasks (Stochastic Binary Networks, Categorical Autoencoders, and Differentiable Logic Gate Networks), Decoupled ST consistently outperforms Identity STE, Softmax STE, and Straight-Through Gumbel-Softmax. Crucially, optimal $(τ_f, τ_b)$ configurations lie far off the diagonal $τ_f = τ_b$, confirming that the two concerns do require different answers and that single-temperature methods are fundamentally constrained.
format Preprint
id arxiv_https___arxiv_org_abs_2410_13331
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving Discrete Optimisation Via Decoupled Straight-Through Estimator
Shah, Rushi
Yan, Mingyuan
Mozer, Michael Curtis
Liu, Dianbo
Machine Learning
Artificial Intelligence
The Straight-Through Estimator (STE) is the dominant method for training neural networks with discrete variables, enabling gradient-based optimisation by routing gradients through a differentiable surrogate. However, existing STE variants conflate two fundamentally distinct concerns: forward-pass stochasticity, which controls exploration and latent space utilisation, and backward-pass gradient dispersion i.e how learning signals are distributed across categories. We show that these concerns are qualitatively different and that tying them to a single temperature parameter leaves significant performance gains untapped. We propose Decoupled Straight-Through (Decoupled ST), a minimal modification that introduces separate temperatures for the forward pass ($τ_f$) and the backward pass ($τ_b$). This simple change enables independent tuning of exploration and gradient dispersion. Across three diverse tasks (Stochastic Binary Networks, Categorical Autoencoders, and Differentiable Logic Gate Networks), Decoupled ST consistently outperforms Identity STE, Softmax STE, and Straight-Through Gumbel-Softmax. Crucially, optimal $(τ_f, τ_b)$ configurations lie far off the diagonal $τ_f = τ_b$, confirming that the two concerns do require different answers and that single-temperature methods are fundamentally constrained.
title Improving Discrete Optimisation Via Decoupled Straight-Through Estimator
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2410.13331