How Does Preconditioning Guide Feature Learning in Deep Neural Networks?
Fuente:
arXiv
Guardado en:
| Autores principales: | Yoshida, Kotaro, Nitanda, Atsushi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Improved Particle Approximation Error for Mean Field Neural Networks
por: Nitanda, Atsushi
Publicado: (2024)
por: Nitanda, Atsushi
Publicado: (2024)
Towards a Unified Analysis of Neural Networks in Nonparametric Instrumental Variable Regression: Optimization and Generalization
por: Chen, Zonghao, et al.
Publicado: (2025)
por: Chen, Zonghao, et al.
Publicado: (2025)
Uniform convergence of the smooth calibration error and its relationship with functional gradient
por: Futami, Futoshi, et al.
Publicado: (2025)
por: Futami, Futoshi, et al.
Publicado: (2025)
Alternating Diffusion for Proximal Sampling with Zeroth Order Queries
por: Takagi, Hirohane, et al.
Publicado: (2026)
por: Takagi, Hirohane, et al.
Publicado: (2026)
Slowly Annealed Langevin Dynamics: Theory and Applications to Training-Free Guided Generation
por: Nitanda, Atsushi, et al.
Publicado: (2026)
por: Nitanda, Atsushi, et al.
Publicado: (2026)
Statistical Analysis of the Sinkhorn Iterations for Two-Sample Schrödinger Bridge Estimation
por: Maeda, Ibuki, et al.
Publicado: (2025)
por: Maeda, Ibuki, et al.
Publicado: (2025)
Mirror Descent Policy Optimisation for Robust Constrained Markov Decision Processes
por: Bossens, David M., et al.
Publicado: (2025)
por: Bossens, David M., et al.
Publicado: (2025)
Direct Distributional Optimization for Provable Alignment of Diffusion Models
por: Kawata, Ryotaro, et al.
Publicado: (2025)
por: Kawata, Ryotaro, et al.
Publicado: (2025)
Why is parameter averaging beneficial in SGD? An objective smoothing perspective
por: Nitanda, Atsushi, et al.
Publicado: (2023)
por: Nitanda, Atsushi, et al.
Publicado: (2023)
Robust Invariant Representation Learning by Distribution Extrapolation
por: Yoshida, Kotaro, et al.
Publicado: (2025)
por: Yoshida, Kotaro, et al.
Publicado: (2025)
How Does Overparameterization Affect Machine Unlearning of Deep Neural Networks?
por: Alon, Gal, et al.
Publicado: (2025)
por: Alon, Gal, et al.
Publicado: (2025)
Intrinsic Wasserstein Rates for Score-Based Generative Models on Smooth Manifolds
por: Fu, Guoji, et al.
Publicado: (2026)
por: Fu, Guoji, et al.
Publicado: (2026)
How Does Gradient Descent Learn Features -- A Local Analysis for Regularized Two-Layer Neural Networks
por: Zhou, Mo, et al.
Publicado: (2024)
por: Zhou, Mo, et al.
Publicado: (2024)
Preconditioning for Physics-Informed Neural Networks
por: Liu, Songming, et al.
Publicado: (2024)
por: Liu, Songming, et al.
Publicado: (2024)
Towards Understanding Variants of Invariant Risk Minimization through the Lens of Calibration
por: Yoshida, Kotaro, et al.
Publicado: (2024)
por: Yoshida, Kotaro, et al.
Publicado: (2024)
On the Neural Feature Ansatz for Deep Neural Networks
por: Tansley, Edward, et al.
Publicado: (2025)
por: Tansley, Edward, et al.
Publicado: (2025)
Koopman-based generalization bound: New aspect for full-rank weights
por: Hashimoto, Yuka, et al.
Publicado: (2023)
por: Hashimoto, Yuka, et al.
Publicado: (2023)
How Does Overparameterization Affect Features?
por: Duzgun, Ahmet Cagri, et al.
Publicado: (2024)
por: Duzgun, Ahmet Cagri, et al.
Publicado: (2024)
Provably Transformers Harness Multi-Concept Word Semantics for Efficient In-Context Learning
por: Bu, Dake, et al.
Publicado: (2024)
por: Bu, Dake, et al.
Publicado: (2024)
Propagation of Chaos for Mean-Field Langevin Dynamics and its Application to Model Ensemble
por: Nitanda, Atsushi, et al.
Publicado: (2025)
por: Nitanda, Atsushi, et al.
Publicado: (2025)
On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning
por: Zhang, Thomas T., et al.
Publicado: (2025)
por: Zhang, Thomas T., et al.
Publicado: (2025)
DP-KFC: Data-Free Preconditioning for Privacy-Preserving Deep Learning
por: Bosch, Marc Molina Van den, et al.
Publicado: (2026)
por: Bosch, Marc Molina Van den, et al.
Publicado: (2026)
Data Diversity as Implicit Regularization: How Does Diversity Shape the Weight Space of Deep Neural Networks?
por: Ba, Yang, et al.
Publicado: (2024)
por: Ba, Yang, et al.
Publicado: (2024)
Preconditioned Inexact Stochastic ADMM for Deep Model
por: Zhou, Shenglong, et al.
Publicado: (2025)
por: Zhou, Shenglong, et al.
Publicado: (2025)
Feature-Guided Analysis of Neural Networks: A Replication Study
por: Formica, Federico, et al.
Publicado: (2025)
por: Formica, Federico, et al.
Publicado: (2025)
Gradient Preconditioning for Efficient and Reliable Reward-Guided Generation
por: Hwang, Jisung, et al.
Publicado: (2026)
por: Hwang, Jisung, et al.
Publicado: (2026)
SPADE: Sparsity-Guided Debugging for Deep Neural Networks
por: Moakhar, Arshia Soltani, et al.
Publicado: (2023)
por: Moakhar, Arshia Soltani, et al.
Publicado: (2023)
How Does a Deep Neural Network Look at Lexical Stress in English Words?
por: Allouche, Itai, et al.
Publicado: (2025)
por: Allouche, Itai, et al.
Publicado: (2025)
Provable Benefit of Curriculum in Transformer Tree-Reasoning Post-Training
por: Bu, Dake, et al.
Publicado: (2025)
por: Bu, Dake, et al.
Publicado: (2025)
Preconditioned Neural Posterior Estimation for Likelihood-free Inference
por: Wang, Xiaoyu, et al.
Publicado: (2024)
por: Wang, Xiaoyu, et al.
Publicado: (2024)
Neighborhood Sampling Does Not Learn the Same Graph Neural Network
por: Niu, Zehao, et al.
Publicado: (2025)
por: Niu, Zehao, et al.
Publicado: (2025)
Slicing Input Features to Accelerate Deep Learning: A Case Study with Graph Neural Networks
por: Xu, Zhengjia, et al.
Publicado: (2024)
por: Xu, Zhengjia, et al.
Publicado: (2024)
A Non-Monotone Preconditioned Trust-Region Method for Neural Network Training
por: Angino, Andrea, et al.
Publicado: (2026)
por: Angino, Andrea, et al.
Publicado: (2026)
On the Stability of Neural Networks in Deep Learning
por: Delattre, Blaise
Publicado: (2025)
por: Delattre, Blaise
Publicado: (2025)
How Does the Lagrangian Guide Safe Reinforcement Learning through Diffusion Models?
por: Cheng, Xiaoyuan, et al.
Publicado: (2026)
por: Cheng, Xiaoyuan, et al.
Publicado: (2026)
Does Your Neural Network Extrapolate? Feature Engineering as Identifiability Bias for OOD Generalization
por: Aguilar, Leonel, et al.
Publicado: (2026)
por: Aguilar, Leonel, et al.
Publicado: (2026)
Inferring Data Preconditions from Deep Learning Models for Trustworthy Prediction in Deployment
por: Ahmed, Shibbir, et al.
Publicado: (2024)
por: Ahmed, Shibbir, et al.
Publicado: (2024)
Data-Parallel Neural Network Training via Nonlinearly Preconditioned Trust-Region Method
por: Alegría, Samuel A. Cruz, et al.
Publicado: (2025)
por: Alegría, Samuel A. Cruz, et al.
Publicado: (2025)
Canonical Regularisation of Wide Feature-Learning Neural Networks
por: Whittle, George, et al.
Publicado: (2026)
por: Whittle, George, et al.
Publicado: (2026)
Learning with Shallow Neural Networks on Cluster-Structured Features
por: Cornacchia, Elisabetta, et al.
Publicado: (2026)
por: Cornacchia, Elisabetta, et al.
Publicado: (2026)
Ejemplares similares
-
Improved Particle Approximation Error for Mean Field Neural Networks
por: Nitanda, Atsushi
Publicado: (2024) -
Towards a Unified Analysis of Neural Networks in Nonparametric Instrumental Variable Regression: Optimization and Generalization
por: Chen, Zonghao, et al.
Publicado: (2025) -
Uniform convergence of the smooth calibration error and its relationship with functional gradient
por: Futami, Futoshi, et al.
Publicado: (2025) -
Alternating Diffusion for Proximal Sampling with Zeroth Order Queries
por: Takagi, Hirohane, et al.
Publicado: (2026) -
Slowly Annealed Langevin Dynamics: Theory and Applications to Training-Free Guided Generation
por: Nitanda, Atsushi, et al.
Publicado: (2026)