DASH: Warm-Starting Neural Network Training in Stationary Settings without Loss of Plasticity
Fuente:
arXiv
Saved in:
| Main Authors: | Shin, Baekrok, Oh, Junsoo, Cho, Hanseul, Yun, Chulhee |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Lightweight Dataset Pruning without Full Training via Example Difficulty and Prediction Uncertainty
by: Cho, Yeseul, et al.
Published: (2025)
by: Cho, Yeseul, et al.
Published: (2025)
Implicit Bias and Loss of Plasticity in Matrix Completion: Depth Promotes Low-Rankness
by: Shin, Baekrok, et al.
Published: (2026)
by: Shin, Baekrok, et al.
Published: (2026)
Provable Benefit of Cutout and CutMix for Feature Learning
by: Oh, Junsoo, et al.
Published: (2024)
by: Oh, Junsoo, et al.
Published: (2024)
Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count
by: Cho, Hanseul, et al.
Published: (2024)
by: Cho, Hanseul, et al.
Published: (2024)
Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure
by: Cho, Hanseul, et al.
Published: (2024)
by: Cho, Hanseul, et al.
Published: (2024)
Uniform Spectral Growth and Convergence of Muon in LoRA-Style Matrix Factorization
by: Kang, Changmin, et al.
Published: (2026)
by: Kang, Changmin, et al.
Published: (2026)
Fundamental Benefit of Alternating Updates in Minimax Optimization
by: Lee, Jaewook, et al.
Published: (2024)
by: Lee, Jaewook, et al.
Published: (2024)
Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification
by: Jung, Hyunji, et al.
Published: (2025)
by: Jung, Hyunji, et al.
Published: (2025)
From Linear to Nonlinear: Provable Weak-to-Strong Generalization through Feature Learning
by: Oh, Junsoo, et al.
Published: (2025)
by: Oh, Junsoo, et al.
Published: (2025)
Neural Network Plasticity and Loss Sharpness
by: Koster, Max, et al.
Published: (2024)
by: Koster, Max, et al.
Published: (2024)
Activation by Interval-wise Dropout: A Simple Way to Prevent Neural Networks from Plasticity Loss
by: Park, Sangyeon, et al.
Published: (2025)
by: Park, Sangyeon, et al.
Published: (2025)
Minor First, Major Last: A Depth-Induced Implicit Bias of Sharpness-Aware Minimization
by: Moon, Chaewon, et al.
Published: (2026)
by: Moon, Chaewon, et al.
Published: (2026)
The Cost of Robustness: Tighter Bounds on Parameter Complexity for Robust Memorization in ReLU Nets
by: Kim, Yujun, et al.
Published: (2025)
by: Kim, Yujun, et al.
Published: (2025)
Fast Training of Recurrent Neural Networks with Stationary State Feedbacks
by: Caillon, Paul, et al.
Published: (2025)
by: Caillon, Paul, et al.
Published: (2025)
Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training
by: Song, Minhak, et al.
Published: (2025)
by: Song, Minhak, et al.
Published: (2025)
Z-Error Loss for Training Neural Networks
by: Godin, Guillaume
Published: (2025)
by: Godin, Guillaume
Published: (2025)
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
by: Kim, Jihwan, et al.
Published: (2026)
by: Kim, Jihwan, et al.
Published: (2026)
Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch Regime
by: Baek, Beomhan, et al.
Published: (2025)
by: Baek, Beomhan, et al.
Published: (2025)
Explaining Neural Networks without Access to Training Data
by: Marton, Sascha, et al.
Published: (2022)
by: Marton, Sascha, et al.
Published: (2022)
Parameter Expanded Stochastic Gradient Markov Chain Monte Carlo
by: Kim, Hyunsu, et al.
Published: (2025)
by: Kim, Hyunsu, et al.
Published: (2025)
Deep Learning Warm Starts for Trajectory Optimization on the International Space Station
by: Banerjee, Somrita, et al.
Published: (2025)
by: Banerjee, Somrita, et al.
Published: (2025)
Non-Stationary Learning of Neural Networks with Automatic Soft Parameter Reset
by: Galashov, Alexandre, et al.
Published: (2024)
by: Galashov, Alexandre, et al.
Published: (2024)
SCPL: Enhancing Neural Network Training Throughput with Decoupled Local Losses and Model Parallelism
by: Ho, Ming-Yao, et al.
Published: (2026)
by: Ho, Ming-Yao, et al.
Published: (2026)
Description of the Training Process of Neural Networks via Ergodic Theorem : Ghost nodes
by: Park, Eun-Ji, et al.
Published: (2025)
by: Park, Eun-Ji, et al.
Published: (2025)
Self-Abstraction Learning for Effective and Stable Training of Deep Neural Networks
by: Cho, Wonyong, et al.
Published: (2026)
by: Cho, Wonyong, et al.
Published: (2026)
Training Greedy Policy for Proposal Batch Selection in Expensive Multi-Objective Combinatorial Optimization
by: Lee, Deokjae, et al.
Published: (2024)
by: Lee, Deokjae, et al.
Published: (2024)
Addressing Loss of Plasticity and Catastrophic Forgetting in Continual Learning
by: Elsayed, Mohamed, et al.
Published: (2024)
by: Elsayed, Mohamed, et al.
Published: (2024)
Plasticity Loss in Deep Reinforcement Learning: A Survey
by: Klein, Timo, et al.
Published: (2024)
by: Klein, Timo, et al.
Published: (2024)
Super Level Sets and Exponential Decay: A Synergistic Approach to Stable Neural Network Training
by: Chaudhary, Jatin, et al.
Published: (2024)
by: Chaudhary, Jatin, et al.
Published: (2024)
Slow and Steady Wins the Race: Maintaining Plasticity with Hare and Tortoise Networks
by: Lee, Hojoon, et al.
Published: (2024)
by: Lee, Hojoon, et al.
Published: (2024)
Do Neural Networks Lose Plasticity in a Gradually Changing World?
by: Liu, Tianhui, et al.
Published: (2026)
by: Liu, Tianhui, et al.
Published: (2026)
DASH: Detection and Assessment of Systematic Hallucinations of VLMs
by: Augustin, Maximilian, et al.
Published: (2025)
by: Augustin, Maximilian, et al.
Published: (2025)
A Study of Plasticity Loss in On-Policy Deep Reinforcement Learning
by: Juliani, Arthur, et al.
Published: (2024)
by: Juliani, Arthur, et al.
Published: (2024)
Barriers for Learning in an Evolving World: Mathematical Understanding of Loss of Plasticity
by: Joudaki, Amir, et al.
Published: (2025)
by: Joudaki, Amir, et al.
Published: (2025)
Mitigating Plasticity Loss in Continual Reinforcement Learning by Reducing Churn
by: Tang, Hongyao, et al.
Published: (2025)
by: Tang, Hongyao, et al.
Published: (2025)
Non-convolutional Graph Neural Networks
by: Wang, Yuanqing, et al.
Published: (2024)
by: Wang, Yuanqing, et al.
Published: (2024)
A Swap-Adversarial Framework for Improving Domain Generalization in Electroencephalography-Based Parkinson's Disease Prediction
by: Jin, Seongwon, et al.
Published: (2026)
by: Jin, Seongwon, et al.
Published: (2026)
CAMEL-CLIP: Channel-aware Multimodal Electroencephalography-text Alignment for Generalizable Brain Foundation Models
by: Choi, Hanseul, et al.
Published: (2026)
by: Choi, Hanseul, et al.
Published: (2026)
Revisiting Clustering of Neural Bandits: Selective Reinitialization for Mitigating Loss of Plasticity
by: Su, Zhiyuan, et al.
Published: (2025)
by: Su, Zhiyuan, et al.
Published: (2025)
Neural Network-based Vehicular Channel Estimation Performance: Effect of Noise in the Training Set
by: Ngorima, Simbarashe Aldrin, et al.
Published: (2025)
by: Ngorima, Simbarashe Aldrin, et al.
Published: (2025)
Similar Items
-
Lightweight Dataset Pruning without Full Training via Example Difficulty and Prediction Uncertainty
by: Cho, Yeseul, et al.
Published: (2025) -
Implicit Bias and Loss of Plasticity in Matrix Completion: Depth Promotes Low-Rankness
by: Shin, Baekrok, et al.
Published: (2026) -
Provable Benefit of Cutout and CutMix for Feature Learning
by: Oh, Junsoo, et al.
Published: (2024) -
Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count
by: Cho, Hanseul, et al.
Published: (2024) -
Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure
by: Cho, Hanseul, et al.
Published: (2024)