Takeuchi's Information Criteria as Generalization Measures for DNNs Close to NTK Regime
Fuente:
arXiv
Guardado en:
| Autores principales: | Naganuma, Hiroki, Suzuki, Taiji, Yokota, Rio, Nomura, Masahiro, Ishikawa, Kohta, Sato, Ikuro |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Masked Gated Linear Unit
por: Tajima, Yukito, et al.
Publicado: (2025)
por: Tajima, Yukito, et al.
Publicado: (2025)
Theoretical Analysis of Robust Overfitting for Wide DNNs: An NTK Approach
por: Fu, Shaopeng, et al.
Publicado: (2023)
por: Fu, Shaopeng, et al.
Publicado: (2023)
Label-NTK Alignments and A Tighter Convergence Bound in the NTK Regime
por: Marreddy, Ruchirinkil, et al.
Publicado: (2026)
por: Marreddy, Ruchirinkil, et al.
Publicado: (2026)
Transformers as Measure-Theoretic Associative Memory: A Statistical Perspective and Minimax Optimality
por: Kawata, Ryotaro, et al.
Publicado: (2026)
por: Kawata, Ryotaro, et al.
Publicado: (2026)
Convergence Bound and Critical Batch Size of Muon Optimizer
por: Sato, Naoki, et al.
Publicado: (2025)
por: Sato, Naoki, et al.
Publicado: (2025)
Local Loss Optimization in the Infinite Width: Stable Parameterization of Predictive Coding Networks and Target Propagation
por: Ishikawa, Satoki, et al.
Publicado: (2024)
por: Ishikawa, Satoki, et al.
Publicado: (2024)
The Mechanism of Weak-to-Strong Generalization: Feature Elicitation from Latent Knowledge
por: Awano, Ryoya, et al.
Publicado: (2026)
por: Awano, Ryoya, et al.
Publicado: (2026)
In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-Learning
por: Wakayama, Tomoya, et al.
Publicado: (2025)
por: Wakayama, Tomoya, et al.
Publicado: (2025)
Training NTK to Generalize with KARE
por: Schwab, Johannes, et al.
Publicado: (2025)
por: Schwab, Johannes, et al.
Publicado: (2025)
Understanding Linear Probing then Fine-tuning Language Models from NTK Perspective
por: Tomihari, Akiyoshi, et al.
Publicado: (2024)
por: Tomihari, Akiyoshi, et al.
Publicado: (2024)
How Does Label Noise Gradient Descent Improve Generalization in the Low SNR Regime?
por: Huang, Wei, et al.
Publicado: (2025)
por: Huang, Wei, et al.
Publicado: (2025)
LoRA Training in the NTK Regime has No Spurious Local Minima
por: Jang, Uijeong, et al.
Publicado: (2024)
por: Jang, Uijeong, et al.
Publicado: (2024)
Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape
por: Kim, Juno, et al.
Publicado: (2024)
por: Kim, Juno, et al.
Publicado: (2024)
Deep Two-Way Matrix Reordering for Relational Data Analysis
por: Watanabe, Chihiro, et al.
Publicado: (2021)
por: Watanabe, Chihiro, et al.
Publicado: (2021)
State Space Models are Provably Comparable to Transformers in Dynamic Token Selection
por: Nishikawa, Naoki, et al.
Publicado: (2024)
por: Nishikawa, Naoki, et al.
Publicado: (2024)
Transformers Provably Solve Parity Efficiently with Chain of Thought
por: Kim, Juno, et al.
Publicado: (2024)
por: Kim, Juno, et al.
Publicado: (2024)
Mean-field Analysis on Two-layer Neural Networks from a Kernel Perspective
por: Takakura, Shokichi, et al.
Publicado: (2024)
por: Takakura, Shokichi, et al.
Publicado: (2024)
AutoLL: Automatic Linear Layout of Graphs based on Deep Neural Network
por: Watanabe, Chihiro, et al.
Publicado: (2021)
por: Watanabe, Chihiro, et al.
Publicado: (2021)
Test time training enhances in-context learning of nonlinear functions
por: Kuwataka, Kento, et al.
Publicado: (2025)
por: Kuwataka, Kento, et al.
Publicado: (2025)
Approximation and Estimation Ability of Transformers for Sequence-to-Sequence Functions with Infinite Dimensional Input
por: Takakura, Shokichi, et al.
Publicado: (2023)
por: Takakura, Shokichi, et al.
Publicado: (2023)
Towards Understanding Variants of Invariant Risk Minimization through the Lens of Calibration
por: Yoshida, Kotaro, et al.
Publicado: (2024)
por: Yoshida, Kotaro, et al.
Publicado: (2024)
Convergence Error Analysis of Reflected Gradient Langevin Dynamics for Globally Optimizing Non-Convex Constrained Problems
por: Sato, Kanji, et al.
Publicado: (2022)
por: Sato, Kanji, et al.
Publicado: (2022)
Koopman-based generalization bound: New aspect for full-rank weights
por: Hashimoto, Yuka, et al.
Publicado: (2023)
por: Hashimoto, Yuka, et al.
Publicado: (2023)
Direct Density Ratio Optimization: A Statistically Consistent Approach to Aligning Large Language Models
por: Higuchi, Rei, et al.
Publicado: (2025)
por: Higuchi, Rei, et al.
Publicado: (2025)
Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
por: Nakamura, Taishi, et al.
Publicado: (2025)
por: Nakamura, Taishi, et al.
Publicado: (2025)
Geometric Insights into Focal Loss: Reducing Curvature for Enhanced Model Calibration
por: Kimura, Masanari, et al.
Publicado: (2024)
por: Kimura, Masanari, et al.
Publicado: (2024)
Towards a Unified Analysis of Neural Networks in Nonparametric Instrumental Variable Regression: Optimization and Generalization
por: Chen, Zonghao, et al.
Publicado: (2025)
por: Chen, Zonghao, et al.
Publicado: (2025)
Hyperparameter Optimization Can Even be Harmful in Off-Policy Learning and How to Deal with It
por: Saito, Yuta, et al.
Publicado: (2024)
por: Saito, Yuta, et al.
Publicado: (2024)
Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency
por: Nishikawa, Naoki, et al.
Publicado: (2025)
por: Nishikawa, Naoki, et al.
Publicado: (2025)
Transformers are Minimax Optimal Nonparametric In-Context Learners
por: Kim, Juno, et al.
Publicado: (2024)
por: Kim, Juno, et al.
Publicado: (2024)
Mamba Can Learn Low-Dimensional Targets In-Context via Test-Time Feature Learning
por: Oh, Junsoo, et al.
Publicado: (2025)
por: Oh, Junsoo, et al.
Publicado: (2025)
Nonconvex Latent Optimally Partitioned Block-Sparse Recovery via Log-Sum and Minimax Concave Penalties
por: Furuhashi, Takanobu, et al.
Publicado: (2026)
por: Furuhashi, Takanobu, et al.
Publicado: (2026)
Revisiting Generalization Measures Beyond IID: An Empirical Study under Distributional Shift
por: Nakai, Sora, et al.
Publicado: (2026)
por: Nakai, Sora, et al.
Publicado: (2026)
From Saddle Points Toward Global Minima: A Newton-Type Method on Wasserstein Space
por: Lascu, Razvan-Andrei, et al.
Publicado: (2026)
por: Lascu, Razvan-Andrei, et al.
Publicado: (2026)
Adversarial Robustness of NTK Neural Networks
por: Hou, Yuxuan
Publicado: (2026)
por: Hou, Yuxuan
Publicado: (2026)
Intrinsic Wasserstein Rates for Score-Based Generative Models on Smooth Manifolds
por: Fu, Guoji, et al.
Publicado: (2026)
por: Fu, Guoji, et al.
Publicado: (2026)
Data Uniformity Improves Training Efficiency and More, with a Convergence Framework Beyond the NTK Regime
por: Wang, Yuqing, et al.
Publicado: (2025)
por: Wang, Yuqing, et al.
Publicado: (2025)
MLPs at the EOC: Spectrum of the NTK
por: Terjék, Dávid, et al.
Publicado: (2025)
por: Terjék, Dávid, et al.
Publicado: (2025)
MLPs at the EOC: Concentration of the NTK
por: Terjék, Dávid, et al.
Publicado: (2025)
por: Terjék, Dávid, et al.
Publicado: (2025)
Balancing Speed and Stability: The Trade-offs of FP8 vs. BF16 Training in LLMs
por: Fujii, Kazuki, et al.
Publicado: (2024)
por: Fujii, Kazuki, et al.
Publicado: (2024)
Ejemplares similares
-
Masked Gated Linear Unit
por: Tajima, Yukito, et al.
Publicado: (2025) -
Theoretical Analysis of Robust Overfitting for Wide DNNs: An NTK Approach
por: Fu, Shaopeng, et al.
Publicado: (2023) -
Label-NTK Alignments and A Tighter Convergence Bound in the NTK Regime
por: Marreddy, Ruchirinkil, et al.
Publicado: (2026) -
Transformers as Measure-Theoretic Associative Memory: A Statistical Perspective and Minimax Optimality
por: Kawata, Ryotaro, et al.
Publicado: (2026) -
Convergence Bound and Critical Batch Size of Muon Optimizer
por: Sato, Naoki, et al.
Publicado: (2025)