Saved in:
| Main Authors: | Huang, Wei, Cao, Yuan, Wang, Haonan, Cao, Xin, Suzuki, Taiji |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2306.13926 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AutoLL: Automatic Linear Layout of Graphs based on Deep Neural Network
by: Watanabe, Chihiro, et al.
Published: (2021)
by: Watanabe, Chihiro, et al.
Published: (2021)
Towards a Unified Analysis of Neural Networks in Nonparametric Instrumental Variable Regression: Optimization and Generalization
by: Chen, Zonghao, et al.
Published: (2025)
by: Chen, Zonghao, et al.
Published: (2025)
Mean-field Analysis on Two-layer Neural Networks from a Kernel Perspective
by: Takakura, Shokichi, et al.
Published: (2024)
by: Takakura, Shokichi, et al.
Published: (2024)
On the Comparison between Multi-modal and Single-modal Contrastive Learning
by: Huang, Wei, et al.
Published: (2024)
by: Huang, Wei, et al.
Published: (2024)
Optimizing Genetic Algorithms with Multilayer Perceptron Networks for Enhancing TinyFace Recognition
by: Al-Batah, Mohammad Subhi, et al.
Published: (2025)
by: Al-Batah, Mohammad Subhi, et al.
Published: (2025)
Simple Full-Spectrum Correlated k-Distribution Model based on Multilayer Perceptron
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
Mamba Can Learn Low-Dimensional Targets In-Context via Test-Time Feature Learning
by: Oh, Junsoo, et al.
Published: (2025)
by: Oh, Junsoo, et al.
Published: (2025)
The Mechanism of Weak-to-Strong Generalization: Feature Elicitation from Latent Knowledge
by: Awano, Ryoya, et al.
Published: (2026)
by: Awano, Ryoya, et al.
Published: (2026)
In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-Learning
by: Wakayama, Tomoya, et al.
Published: (2025)
by: Wakayama, Tomoya, et al.
Published: (2025)
Direct Density Ratio Optimization: A Statistically Consistent Approach to Aligning Large Language Models
by: Higuchi, Rei, et al.
Published: (2025)
by: Higuchi, Rei, et al.
Published: (2025)
Latent Space Topology Evolution in Multilayer Perceptrons
by: Paluzo-Hidalgo, Eduardo
Published: (2025)
by: Paluzo-Hidalgo, Eduardo
Published: (2025)
Penalty Learning for Optimal Partitioning using Multilayer Perceptron
by: Nguyen, Tung L, et al.
Published: (2024)
by: Nguyen, Tung L, et al.
Published: (2024)
Peptidomic-Based Prediction Model for Coronary Heart Disease Using a Multilayer Perceptron Neural Network
by: Celis-Porras, Jesus
Published: (2025)
by: Celis-Porras, Jesus
Published: (2025)
On the Optimization and Generalization of Two-layer Transformers with Sign Gradient Descent
by: Li, Bingrui, et al.
Published: (2024)
by: Li, Bingrui, et al.
Published: (2024)
Ranked Set Sampling-Based Multilayer Perceptron: Improving Generalization via Variance-Based Bounds
by: Li, Feijiang, et al.
Published: (2025)
by: Li, Feijiang, et al.
Published: (2025)
Kolmogorov-Arnold Networks in Low-Data Regimes: A Comparative Study with Multilayer Perceptrons
by: Pourkamali-Anaraki, Farhad
Published: (2024)
by: Pourkamali-Anaraki, Farhad
Published: (2024)
Approximation and Estimation Ability of Transformers for Sequence-to-Sequence Functions with Infinite Dimensional Input
by: Takakura, Shokichi, et al.
Published: (2023)
by: Takakura, Shokichi, et al.
Published: (2023)
Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape
by: Kim, Juno, et al.
Published: (2024)
by: Kim, Juno, et al.
Published: (2024)
Deep Two-Way Matrix Reordering for Relational Data Analysis
by: Watanabe, Chihiro, et al.
Published: (2021)
by: Watanabe, Chihiro, et al.
Published: (2021)
State Space Models are Provably Comparable to Transformers in Dynamic Token Selection
by: Nishikawa, Naoki, et al.
Published: (2024)
by: Nishikawa, Naoki, et al.
Published: (2024)
Transformers Provably Solve Parity Efficiently with Chain of Thought
by: Kim, Juno, et al.
Published: (2024)
by: Kim, Juno, et al.
Published: (2024)
Test time training enhances in-context learning of nonlinear functions
by: Kuwataka, Kento, et al.
Published: (2025)
by: Kuwataka, Kento, et al.
Published: (2025)
Transformers as Measure-Theoretic Associative Memory: A Statistical Perspective and Minimax Optimality
by: Kawata, Ryotaro, et al.
Published: (2026)
by: Kawata, Ryotaro, et al.
Published: (2026)
Theoretical Compression Bounds for Wide Multilayer Perceptrons
by: Cheairi, Houssam El, et al.
Published: (2025)
by: Cheairi, Houssam El, et al.
Published: (2025)
Generalization Bound of Gradient Flow through Training Trajectory and Data-dependent Kernel
by: Chen, Yilan, et al.
Published: (2025)
by: Chen, Yilan, et al.
Published: (2025)
Universal Approximation Theorem for Input-Connected Multilayer Perceptrons
by: Ismailov, Vugar
Published: (2026)
by: Ismailov, Vugar
Published: (2026)
Direct Distributional Optimization for Provable Alignment of Diffusion Models
by: Kawata, Ryotaro, et al.
Published: (2025)
by: Kawata, Ryotaro, et al.
Published: (2025)
Bespoke Approximation of Multiplication-Accumulation and Activation Targeting Printed Multilayer Perceptrons
by: Afentaki, Florentia, et al.
Published: (2023)
by: Afentaki, Florentia, et al.
Published: (2023)
How Neural Reward Models Learn Features for Policy Optimization: A Single-Index Analysis
by: Higuchi, Rei, et al.
Published: (2026)
by: Higuchi, Rei, et al.
Published: (2026)
Unveil Benign Overfitting for Transformer in Vision: Training Dynamics, Convergence, and Generalization
by: Jiang, Jiarui, et al.
Published: (2024)
by: Jiang, Jiarui, et al.
Published: (2024)
Batch Matrix-form Equations and Implementation of Multilayer Perceptrons
by: Wesselink, Wieger, et al.
Published: (2025)
by: Wesselink, Wieger, et al.
Published: (2025)
Trained Mamba Emulates Online Gradient Descent in In-Context Linear Regression
by: Jiang, Jiarui, et al.
Published: (2025)
by: Jiang, Jiarui, et al.
Published: (2025)
[Experiments & Analysis] Evaluating the Feasibility of Sampling-Based Techniques for Training Multilayer Perceptrons
by: Ebrahimi, Sana, et al.
Published: (2023)
by: Ebrahimi, Sana, et al.
Published: (2023)
Neural network learns low-dimensional polynomials with SGD near the information-theoretic limit
by: Lee, Jason D., et al.
Published: (2024)
by: Lee, Jason D., et al.
Published: (2024)
Provably Neural Active Learning Succeeds via Prioritizing Perplexing Samples
by: Bu, Dake, et al.
Published: (2024)
by: Bu, Dake, et al.
Published: (2024)
On the Learning Dynamics of Two-layer Linear Networks with Label Noise SGD
by: Zhang, Tongcheng, et al.
Published: (2026)
by: Zhang, Tongcheng, et al.
Published: (2026)
From Saddle Points Toward Global Minima: A Newton-Type Method on Wasserstein Space
by: Lascu, Razvan-Andrei, et al.
Published: (2026)
by: Lascu, Razvan-Andrei, et al.
Published: (2026)
Optimality and Adaptivity of Deep Neural Features for Instrumental Variable Regression
by: Kim, Juno, et al.
Published: (2025)
by: Kim, Juno, et al.
Published: (2025)
Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency
by: Nishikawa, Naoki, et al.
Published: (2025)
by: Nishikawa, Naoki, et al.
Published: (2025)
Transformers are Minimax Optimal Nonparametric In-Context Learners
by: Kim, Juno, et al.
Published: (2024)
by: Kim, Juno, et al.
Published: (2024)
Similar Items
-
AutoLL: Automatic Linear Layout of Graphs based on Deep Neural Network
by: Watanabe, Chihiro, et al.
Published: (2021) -
Towards a Unified Analysis of Neural Networks in Nonparametric Instrumental Variable Regression: Optimization and Generalization
by: Chen, Zonghao, et al.
Published: (2025) -
Mean-field Analysis on Two-layer Neural Networks from a Kernel Perspective
by: Takakura, Shokichi, et al.
Published: (2024) -
On the Comparison between Multi-modal and Single-modal Contrastive Learning
by: Huang, Wei, et al.
Published: (2024) -
Optimizing Genetic Algorithms with Multilayer Perceptron Networks for Enhancing TinyFace Recognition
by: Al-Batah, Mohammad Subhi, et al.
Published: (2025)