Custom Gradient Estimators are Straight-Through Estimators in Disguise
Fuente:
arXiv
Saved in:
| Main Authors: | Schoenbauer, Matt, Moro, Daniele, Lew, Lukasz, Howard, Andrew |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving the Straight-Through Estimator with Zeroth-Order Information
by: Yang, Ningfeng, et al.
Published: (2025)
by: Yang, Ningfeng, et al.
Published: (2025)
PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression
by: Malinovskii, Vladimir, et al.
Published: (2024)
by: Malinovskii, Vladimir, et al.
Published: (2024)
Improving Discrete Optimisation Via Decoupled Straight-Through Estimator
by: Shah, Rushi, et al.
Published: (2024)
by: Shah, Rushi, et al.
Published: (2024)
LLM Safety Alignment is Divergence Estimation in Disguise
by: Haldar, Rajdeep, et al.
Published: (2025)
by: Haldar, Rajdeep, et al.
Published: (2025)
Extending Straight-Through Estimation for Robust Neural Networks on Analog CIM Hardware
by: Feng, Yuannuo, et al.
Published: (2025)
by: Feng, Yuannuo, et al.
Published: (2025)
High-Dimensional Learning Dynamics of Quantized Models with Straight-Through Estimator
by: Ichikawa, Yuma, et al.
Published: (2025)
by: Ichikawa, Yuma, et al.
Published: (2025)
Setting the Record Straight on Transformer Oversmoothing
by: Dovonon, Gbètondji J-S, et al.
Published: (2024)
by: Dovonon, Gbètondji J-S, et al.
Published: (2024)
Convexity in Disguise: A Theoretical Framework for Nonconvex Low-Rank Matrix Estimation
by: Cui, Chengyu, et al.
Published: (2026)
by: Cui, Chengyu, et al.
Published: (2026)
PikeLPN: Mitigating Overlooked Inefficiencies of Low-Precision Neural Networks
by: Neseem, Marina, et al.
Published: (2024)
by: Neseem, Marina, et al.
Published: (2024)
Position: Explainable AI is Causality in Disguise
by: Karimi, Amir-Hossein
Published: (2026)
by: Karimi, Amir-Hossein
Published: (2026)
Refining Covariance Matrix Estimation in Stochastic Gradient Descent Through Bias Reduction
by: Wei, Ziyang, et al.
Published: (2026)
by: Wei, Ziyang, et al.
Published: (2026)
Beyond Discreteness: Sample Complexity Analysis of Straight-Through Estimator for 1-bit Quantization
by: Jeong, Halyun, et al.
Published: (2025)
by: Jeong, Halyun, et al.
Published: (2025)
ABCD: All Biases Come Disguised
by: Nowak, Mateusz, et al.
Published: (2026)
by: Nowak, Mateusz, et al.
Published: (2026)
Semiparametric Efficient Bilevel Gradient Estimation
by: Khoury, Fares El, et al.
Published: (2026)
by: Khoury, Fares El, et al.
Published: (2026)
Gradient Estimation with Discrete Stein Operators
by: Shi, Jiaxin, et al.
Published: (2022)
by: Shi, Jiaxin, et al.
Published: (2022)
On the Convergence and Straightness of Rectified Flow
by: Bansal, Vansh, et al.
Published: (2024)
by: Bansal, Vansh, et al.
Published: (2024)
Catastrophic Overfitting: A Potential Blessing in Disguise
by: Zhao, Mengnan, et al.
Published: (2024)
by: Zhao, Mengnan, et al.
Published: (2024)
Disguised Copyright Infringement of Latent Diffusion Models
by: Lu, Yiwei, et al.
Published: (2024)
by: Lu, Yiwei, et al.
Published: (2024)
Binomial Gradient-Based Meta-Learning for Enhanced Meta-Gradient Estimation
by: Zhang, Yilang, et al.
Published: (2026)
by: Zhang, Yilang, et al.
Published: (2026)
Multiple Importance Sampling for Stochastic Gradient Estimation
by: Salaün, Corentin, et al.
Published: (2024)
by: Salaün, Corentin, et al.
Published: (2024)
Adaptive Feedforward Gradient Estimation in Neural ODEs
by: Dabounou, Jaouad
Published: (2024)
by: Dabounou, Jaouad
Published: (2024)
Generalizing Stochastic Smoothing for Differentiation and Gradient Estimation
by: Petersen, Felix, et al.
Published: (2024)
by: Petersen, Felix, et al.
Published: (2024)
Towards Scalable Backpropagation-Free Gradient Estimation
by: Wang, Daniel, et al.
Published: (2025)
by: Wang, Daniel, et al.
Published: (2025)
A Wiener Process Perspective on Local Intrinsic Dimension Estimation Methods
by: Tempczyk, Piotr, et al.
Published: (2024)
by: Tempczyk, Piotr, et al.
Published: (2024)
Density Ratio Estimation via Sampling along Generalized Geodesics on Statistical Manifolds
by: Kimura, Masanari, et al.
Published: (2024)
by: Kimura, Masanari, et al.
Published: (2024)
Chain-of-Thought Unfaithfulness as Disguised Accuracy
by: Bentham, Oliver, et al.
Published: (2024)
by: Bentham, Oliver, et al.
Published: (2024)
Gumbel-Softmax Flow Matching with Straight-Through Guidance for Controllable Biological Sequence Generation
by: Tang, Sophia, et al.
Published: (2025)
by: Tang, Sophia, et al.
Published: (2025)
Differentiable Cost-Parameterized Monge Map Estimators
by: Howard, Samuel, et al.
Published: (2024)
by: Howard, Samuel, et al.
Published: (2024)
Customer Lifetime Value Prediction with Uncertainty Estimation Using Monte Carlo Dropout
by: Cao, Xinzhe, et al.
Published: (2024)
by: Cao, Xinzhe, et al.
Published: (2024)
An Eulerian Perspective on Straight-Line Sampling
by: Tsimpos, Panos, et al.
Published: (2025)
by: Tsimpos, Panos, et al.
Published: (2025)
Parallel Momentum Methods Under Biased Gradient Estimations
by: Beikmohammadi, Ali, et al.
Published: (2024)
by: Beikmohammadi, Ali, et al.
Published: (2024)
Fast and Unified Path Gradient Estimators for Normalizing Flows
by: Vaitl, Lorenz, et al.
Published: (2024)
by: Vaitl, Lorenz, et al.
Published: (2024)
Entropy-regularized Gradient Estimators for Approximate Bayesian Inference
by: Kaur, Jasmeet
Published: (2025)
by: Kaur, Jasmeet
Published: (2025)
Rao-Blackwell Gradient Estimators for Equivariant Denoising Diffusion
by: Tong, Vinh, et al.
Published: (2025)
by: Tong, Vinh, et al.
Published: (2025)
Generalized Advantage Estimation for Distributional Policy Gradients
by: Shaik, Shahil, et al.
Published: (2025)
by: Shaik, Shahil, et al.
Published: (2025)
Straight-Through meets Sparse Recovery: the Support Exploration Algorithm
by: Mohamed, Mimoun, et al.
Published: (2023)
by: Mohamed, Mimoun, et al.
Published: (2023)
Bounding Evidence and Estimating Log-Likelihood in VAE
by: Struski, Łukasz, et al.
Published: (2022)
by: Struski, Łukasz, et al.
Published: (2022)
Multi-scale Autoregressive Models are Laplacian, Discrete, and Latent Diffusion Models in Disguise
by: Hong, Steve, et al.
Published: (2025)
by: Hong, Steve, et al.
Published: (2025)
Learning Straight Flows by Learning Curved Interpolants
by: Shankar, Shiv, et al.
Published: (2025)
by: Shankar, Shiv, et al.
Published: (2025)
Leveraging Gradients for Unsupervised Accuracy Estimation under Distribution Shift
by: Xie, Renchunzi, et al.
Published: (2024)
by: Xie, Renchunzi, et al.
Published: (2024)
Similar Items
-
Improving the Straight-Through Estimator with Zeroth-Order Information
by: Yang, Ningfeng, et al.
Published: (2025) -
PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression
by: Malinovskii, Vladimir, et al.
Published: (2024) -
Improving Discrete Optimisation Via Decoupled Straight-Through Estimator
by: Shah, Rushi, et al.
Published: (2024) -
LLM Safety Alignment is Divergence Estimation in Disguise
by: Haldar, Rajdeep, et al.
Published: (2025) -
Extending Straight-Through Estimation for Robust Neural Networks on Analog CIM Hardware
by: Feng, Yuannuo, et al.
Published: (2025)