Accelerated Training through Iterative Gradient Propagation Along the Residual Path
Fuente:
arXiv
Saved in:
| Main Authors: | Fagnou, Erwan, Caillon, Paul, Delattre, Blaise, Allauzen, Alexandre |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Trading Complexity for Expressivity Through Structured Generalized Linear Token Mixing
by: Fagnou, Erwan, et al.
Published: (2026)
by: Fagnou, Erwan, et al.
Published: (2026)
Chain and Causal Attention for Efficient Entity Tracking
by: Fagnou, Erwan, et al.
Published: (2024)
by: Fagnou, Erwan, et al.
Published: (2024)
Bridging the Theoretical Gap in Randomized Smoothing
by: Delattre, Blaise, et al.
Published: (2025)
by: Delattre, Blaise, et al.
Published: (2025)
Forward Only Learning for Orthogonal Neural Networks of any Depth
by: Caillon, Paul, et al.
Published: (2025)
by: Caillon, Paul, et al.
Published: (2025)
Fast Training of Recurrent Neural Networks with Stationary State Feedbacks
by: Caillon, Paul, et al.
Published: (2025)
by: Caillon, Paul, et al.
Published: (2025)
Structured-Sparse Attention for Entity Tracking with Subquadratic Sequence Complexity
by: Zhao, Hangyue, et al.
Published: (2026)
by: Zhao, Hangyue, et al.
Published: (2026)
Conditional Distribution Quantization in Machine Learning
by: Delattre, Blaise, et al.
Published: (2025)
by: Delattre, Blaise, et al.
Published: (2025)
Spectral Norm of Convolutional Layers with Circular and Zero Paddings
by: Delattre, Blaise, et al.
Published: (2024)
by: Delattre, Blaise, et al.
Published: (2024)
The Lipschitz-Variance-Margin Tradeoff for Enhanced Randomized Smoothing
by: Delattre, Blaise, et al.
Published: (2023)
by: Delattre, Blaise, et al.
Published: (2023)
Limits of Resolution Equivariance in Fourier Neural Operators
by: Colagrande, Alex, et al.
Published: (2026)
by: Colagrande, Alex, et al.
Published: (2026)
On the Stability of Neural Networks in Deep Learning
by: Delattre, Blaise
Published: (2025)
by: Delattre, Blaise
Published: (2025)
Certified Robustness under Heterogeneous Perturbations via Hybrid Randomized Smoothing
by: Delattre, Blaise, et al.
Published: (2026)
by: Delattre, Blaise, et al.
Published: (2026)
Linear Attention with Global Context: A Multipole Attention Mechanism for Vision and Physics
by: Colagrande, Alex, et al.
Published: (2025)
by: Colagrande, Alex, et al.
Published: (2025)
Accelerating Neural Network Training Along Sharp and Flat Directions
by: Zakarin, Daniyar, et al.
Published: (2025)
by: Zakarin, Daniyar, et al.
Published: (2025)
Differentially Private Gradient Flow based on the Sliced Wasserstein Distance
by: Sebag, Ilana, et al.
Published: (2023)
by: Sebag, Ilana, et al.
Published: (2023)
Polynomial Mixing for Efficient Self-supervised Speech Encoders
by: Feillet, Eva, et al.
Published: (2026)
by: Feillet, Eva, et al.
Published: (2026)
Exploring Precision and Recall to assess the quality and diversity of LLMs
by: Bronnec, Florian Le, et al.
Published: (2024)
by: Bronnec, Florian Le, et al.
Published: (2024)
GeoDirDock: Guiding Docking Along Geodesic Paths
by: Miñán, Raúl, et al.
Published: (2024)
by: Miñán, Raúl, et al.
Published: (2024)
PRISM: Parallel Residual Iterative Sequence Model
by: Jiang, Jie, et al.
Published: (2026)
by: Jiang, Jie, et al.
Published: (2026)
Understanding Gradient Descent through the Training Jacobian
by: Belrose, Nora, et al.
Published: (2024)
by: Belrose, Nora, et al.
Published: (2024)
Gradient Residual Connections
by: Pan, Yangchen, et al.
Published: (2026)
by: Pan, Yangchen, et al.
Published: (2026)
Improving Diversity in Language Models: When Temperature Fails, Change the Loss
by: Verine, Alexandre, et al.
Published: (2025)
by: Verine, Alexandre, et al.
Published: (2025)
On the MIA Vulnerability Gap Between Private GANs and Diffusion Models
by: Sebag, Ilana, et al.
Published: (2025)
by: Sebag, Ilana, et al.
Published: (2025)
GNNs Meet Sequence Models Along the Shortest-Path: an Expressive Method for Link Prediction
by: Ferrini, Francesco, et al.
Published: (2025)
by: Ferrini, Francesco, et al.
Published: (2025)
Quantum Equilibrium Propagation: Gradient-Descent Training of Quantum Systems
by: Scellier, Benjamin
Published: (2024)
by: Scellier, Benjamin
Published: (2024)
Path-Sampled Integrated Gradients
by: Kamalov, Firuz, et al.
Published: (2026)
by: Kamalov, Firuz, et al.
Published: (2026)
Iterate to Accelerate: A Unified Framework for Iterative Reasoning and Feedback Convergence
by: Fein-Ashley, Jacob
Published: (2025)
by: Fein-Ashley, Jacob
Published: (2025)
QF: Quick Feedforward AI Model Training without Gradient Back Propagation
by: Qi, Feng
Published: (2025)
by: Qi, Feng
Published: (2025)
DIVEBATCH: Accelerating Model Training Through Gradient-Diversity Aware Batch Size Adaptation
by: Chen, Yuen, et al.
Published: (2025)
by: Chen, Yuen, et al.
Published: (2025)
Fast and Effective GNN Training through Sequences of Random Path Graphs
by: Bonchi, Francesco, et al.
Published: (2023)
by: Bonchi, Francesco, et al.
Published: (2023)
Vanilla Gradient Descent for Oblique Decision Trees
by: Panda, Subrat Prasad, et al.
Published: (2024)
by: Panda, Subrat Prasad, et al.
Published: (2024)
Path Gradients after Flow Matching
by: Vaitl, Lorenz, et al.
Published: (2025)
by: Vaitl, Lorenz, et al.
Published: (2025)
Accelerating Electron Dynamics Simulations through Machine Learned Time Propagators
by: Shah, Karan, et al.
Published: (2024)
by: Shah, Karan, et al.
Published: (2024)
Residual-as-Teacher: Mitigating Bias Propagation in Student--Teacher Estimation
by: Yamamoto, Kakei, et al.
Published: (2026)
by: Yamamoto, Kakei, et al.
Published: (2026)
Stacking as Accelerated Gradient Descent
by: Agarwal, Naman, et al.
Published: (2024)
by: Agarwal, Naman, et al.
Published: (2024)
Gradient Iterated Temporal-Difference Learning
by: Vincent, Théo, et al.
Published: (2026)
by: Vincent, Théo, et al.
Published: (2026)
The Phase Is the Gradient: Equilibrium Propagation for Frequency Learning in Kuramoto Networks
by: Ahmadi, Mani Rash
Published: (2026)
by: Ahmadi, Mani Rash
Published: (2026)
Dual Path Attribution: Efficient Attribution for SwiGLU-Transformers through Layer-Wise Target Propagation
by: Jantsch, Lasse Marten, et al.
Published: (2026)
by: Jantsch, Lasse Marten, et al.
Published: (2026)
Latent Optimal Paths by Gumbel Propagation for Variational Bayesian Dynamic Programming
by: Niu, Xinlei, et al.
Published: (2023)
by: Niu, Xinlei, et al.
Published: (2023)
Accelerating LLM Pre-Training through Flat-Direction Dynamics Enhancement
by: Zhu, Shuchen, et al.
Published: (2026)
by: Zhu, Shuchen, et al.
Published: (2026)
Similar Items
-
Trading Complexity for Expressivity Through Structured Generalized Linear Token Mixing
by: Fagnou, Erwan, et al.
Published: (2026) -
Chain and Causal Attention for Efficient Entity Tracking
by: Fagnou, Erwan, et al.
Published: (2024) -
Bridging the Theoretical Gap in Randomized Smoothing
by: Delattre, Blaise, et al.
Published: (2025) -
Forward Only Learning for Orthogonal Neural Networks of any Depth
by: Caillon, Paul, et al.
Published: (2025) -
Fast Training of Recurrent Neural Networks with Stationary State Feedbacks
by: Caillon, Paul, et al.
Published: (2025)