Saved in:
| Main Author: | Park, Jin Hyun |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2201.11653 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adaptive multiple optimal learning factors for neural network training
by: Challagundla, Jeshwanth
Published: (2024)
by: Challagundla, Jeshwanth
Published: (2024)
Anon: Extrapolating Adaptivity Beyond SGD and Adam
by: Zhang, Yiheng, et al.
Published: (2026)
by: Zhang, Yiheng, et al.
Published: (2026)
Flow reconstruction in time-varying geometries using graph neural networks
by: Danciu, Bogdan A., et al.
Published: (2024)
by: Danciu, Bogdan A., et al.
Published: (2024)
LayerCollapse: Adaptive compression of neural networks
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2023)
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2023)
Conditional computation in neural networks: principles and research trends
by: Scardapane, Simone, et al.
Published: (2024)
by: Scardapane, Simone, et al.
Published: (2024)
Understanding the learned look-ahead behavior of chess neural networks
by: Cruz, Diogo
Published: (2025)
by: Cruz, Diogo
Published: (2025)
Time-varying Interaction Graph ODE for Dynamic Graph Representation Learning
by: Wang, Xiaoyi, et al.
Published: (2026)
by: Wang, Xiaoyi, et al.
Published: (2026)
Discrete Dictionary-based Decomposition Layer for Structured Representation Learning
by: Park, Taewon, et al.
Published: (2024)
by: Park, Taewon, et al.
Published: (2024)
Enhancing DP-SGD through Non-monotonous Adaptive Scaling Gradient Weight
by: Huang, Tao, et al.
Published: (2024)
by: Huang, Tao, et al.
Published: (2024)
Adaptive Selection of LoRA Components in Privacy-Preserving Federated Learning
by: Kim, Myoungjun, et al.
Published: (2026)
by: Kim, Myoungjun, et al.
Published: (2026)
Why Adam Can Beat SGD: Second-Moment Normalization Yields Sharper Tails
by: Jin, Ruinan, et al.
Published: (2026)
by: Jin, Ruinan, et al.
Published: (2026)
Less is more: Embracing sparsity and interpolation with Esiformer for time series forecasting
by: Guo, Yangyang, et al.
Published: (2024)
by: Guo, Yangyang, et al.
Published: (2024)
Interpolating neural network: A novel unification of machine learning and interpolation theory
by: Park, Chanwook, et al.
Published: (2024)
by: Park, Chanwook, et al.
Published: (2024)
TLDR: Unsupervised Goal-Conditioned RL via Temporal Distance-Aware Representations
by: Bae, Junik, et al.
Published: (2024)
by: Bae, Junik, et al.
Published: (2024)
Multiverse: Language-Conditioned Multi-Game Level Blending via Shared Representation
by: Baek, In-Chang, et al.
Published: (2026)
by: Baek, In-Chang, et al.
Published: (2026)
PIG: Physics-Informed Gaussians as Adaptive Parametric Mesh Representations
by: Kang, Namgyu, et al.
Published: (2024)
by: Kang, Namgyu, et al.
Published: (2024)
MORE-CLEAR: Multimodal Offline Reinforcement learning for Clinical notes Leveraged Enhanced State Representation
by: Lim, Yooseok, et al.
Published: (2025)
by: Lim, Yooseok, et al.
Published: (2025)
Bootstrap SGD: Algorithmic Stability and Robustness
by: Christmann, Andreas, et al.
Published: (2024)
by: Christmann, Andreas, et al.
Published: (2024)
MoE-PHDS: One MoE checkpoint for flexible runtime sparsity
by: Hannah, Lauren. A, et al.
Published: (2025)
by: Hannah, Lauren. A, et al.
Published: (2025)
Feature learning as alignment: a structural property of gradient descent in non-linear neural networks
by: Beaglehole, Daniel, et al.
Published: (2024)
by: Beaglehole, Daniel, et al.
Published: (2024)
On permutation-invariant neural networks
by: Kimura, Masanari, et al.
Published: (2024)
by: Kimura, Masanari, et al.
Published: (2024)
Sobolev acceleration for neural networks
by: Oh, Jong Kwon, et al.
Published: (2025)
by: Oh, Jong Kwon, et al.
Published: (2025)
Attention mechanisms in neural networks
by: Hays, Hasi
Published: (2026)
by: Hays, Hasi
Published: (2026)
APOLLO: SGD-like Memory, AdamW-level Performance
by: Zhu, Hanqing, et al.
Published: (2024)
by: Zhu, Hanqing, et al.
Published: (2024)
Target noise: A pre-training based neural network initialization for efficient high resolution learning
by: Wang, Shaowen, et al.
Published: (2026)
by: Wang, Shaowen, et al.
Published: (2026)
Zero-Incentive Dynamics: a look at reward sparsity through the lens of unrewarded subgoals
by: Molinghen, Yannick, et al.
Published: (2025)
by: Molinghen, Yannick, et al.
Published: (2025)
Accumulative SGD Influence Estimation for Data Attribution
by: Shi, Yunxiao, et al.
Published: (2025)
by: Shi, Yunxiao, et al.
Published: (2025)
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
by: Kim, Jihwan, et al.
Published: (2026)
by: Kim, Jihwan, et al.
Published: (2026)
Principles of Lipschitz continuity in neural networks
by: Luo, Róisín
Published: (2026)
by: Luo, Róisín
Published: (2026)
Linearity-based neural network compression
by: Dobler, Silas, et al.
Published: (2025)
by: Dobler, Silas, et al.
Published: (2025)
Dynamic sparsity in tree-structured feed-forward layers at scale
by: Sedghi, Reza, et al.
Published: (2026)
by: Sedghi, Reza, et al.
Published: (2026)
On-site estimation of battery electrochemical parameters via transfer learning based physics-informed neural network approach
by: Yeregui, Josu, et al.
Published: (2025)
by: Yeregui, Josu, et al.
Published: (2025)
RQP-SGD: Differential Private Machine Learning through Noisy SGD and Randomized Quantization
by: Feng, Ce, et al.
Published: (2024)
by: Feng, Ce, et al.
Published: (2024)
Diagonalisation SGD: Fast & Convergent SGD for Non-Differentiable Models via Reparameterisation and Smoothing
by: Wagner, Dominik, et al.
Published: (2024)
by: Wagner, Dominik, et al.
Published: (2024)
From 2:4 to 8:16 sparsity patterns in LLMs for Outliers and Weights with Variance Correction
by: Maximov, Egor, et al.
Published: (2025)
by: Maximov, Egor, et al.
Published: (2025)
Beyond 2:4: exploring V:N:M sparsity for efficient transformer inference on GPUs
by: Zhao, Kang, et al.
Published: (2024)
by: Zhao, Kang, et al.
Published: (2024)
Worker Disagreement Reveals Sharp Directions in Local SGD
by: Dimlioglu, Tolga, et al.
Published: (2026)
by: Dimlioglu, Tolga, et al.
Published: (2026)
Similarity-based context aware continual learning for spiking neural networks
by: Han, Bing, et al.
Published: (2024)
by: Han, Bing, et al.
Published: (2024)
Applying graph neural network to SupplyGraph for supply chain network
by: Han, Kihwan
Published: (2024)
by: Han, Kihwan
Published: (2024)
Active teacher selection for reward learning
by: Freedman, Rachel, et al.
Published: (2023)
by: Freedman, Rachel, et al.
Published: (2023)
Similar Items
-
Adaptive multiple optimal learning factors for neural network training
by: Challagundla, Jeshwanth
Published: (2024) -
Anon: Extrapolating Adaptivity Beyond SGD and Adam
by: Zhang, Yiheng, et al.
Published: (2026) -
Flow reconstruction in time-varying geometries using graph neural networks
by: Danciu, Bogdan A., et al.
Published: (2024) -
LayerCollapse: Adaptive compression of neural networks
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2023) -
Conditional computation in neural networks: principles and research trends
by: Scardapane, Simone, et al.
Published: (2024)