Revisiting Softmax Masking: Stop Gradient for Enhancing Stability in Replay-based Continual Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Hoyong, Kwon, Minchan, Kim, Kangil |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Asymptotic Midpoint Mixup for Margin Balancing and Moderate Broadening
by: Kim, Hoyong, et al.
Published: (2024)
by: Kim, Hoyong, et al.
Published: (2024)
Multiple Invertible and Partial-Equivariant Function for Latent Vector Transformation to Enhance Disentanglement in VAEs
by: Jung, Hee-Jun, et al.
Published: (2025)
by: Jung, Hee-Jun, et al.
Published: (2025)
CFASL: Composite Factor-Aligned Symmetry Learning for Disentanglement in Variational AutoEncoder
by: Jung, Hee-Jun, et al.
Published: (2024)
by: Jung, Hee-Jun, et al.
Published: (2024)
Teaching AI to Remember: Insights from Brain-Inspired Replay in Continual Learning
by: Kim, Jina
Published: (2025)
by: Kim, Jina
Published: (2025)
Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models
by: Abbes, Istabrak, et al.
Published: (2025)
by: Abbes, Istabrak, et al.
Published: (2025)
Efficient Parallel Audio Generation using Group Masked Language Modeling
by: Jeong, Myeonghun, et al.
Published: (2024)
by: Jeong, Myeonghun, et al.
Published: (2024)
Learning from Matured Dumb Teacher for Fine Generalization
by: Jung, HeeSeung, et al.
Published: (2021)
by: Jung, HeeSeung, et al.
Published: (2021)
Adaptive Sparsified Graph Learning Framework for Vessel Behavior Anomalies
by: Kim, Jeehong, et al.
Published: (2025)
by: Kim, Jeehong, et al.
Published: (2025)
Logit Dynamics in Softmax Policy Gradient Methods
by: Li, Yingru
Published: (2025)
by: Li, Yingru
Published: (2025)
FedSOL: Stabilized Orthogonal Learning with Proximal Restrictions in Federated Learning
by: Lee, Gihun, et al.
Published: (2023)
by: Lee, Gihun, et al.
Published: (2023)
Spatio-Temporal Graphs Beyond Grids: Benchmark for Maritime Anomaly Detection
by: Kim, Jeehong, et al.
Published: (2025)
by: Kim, Jeehong, et al.
Published: (2025)
Better Generative Replay for Continual Federated Learning
by: Qi, Daiqing, et al.
Published: (2023)
by: Qi, Daiqing, et al.
Published: (2023)
Feature Structure Distillation with Centered Kernel Alignment in BERT Transferring
by: Jung, Hee-Jun, et al.
Published: (2022)
by: Jung, Hee-Jun, et al.
Published: (2022)
Continual Offline Reinforcement Learning via Diffusion-based Dual Generative Replay
by: Liu, Jinmei, et al.
Published: (2024)
by: Liu, Jinmei, et al.
Published: (2024)
Dual-LoRA and Quality-Enhanced Pseudo Replay for Multimodal Continual Food Learning
by: Wu, Xinlan, et al.
Published: (2025)
by: Wu, Xinlan, et al.
Published: (2025)
DSLR: Diversity Enhancement and Structure Learning for Rehearsal-based Graph Continual Learning
by: Choi, Seungyoon, et al.
Published: (2024)
by: Choi, Seungyoon, et al.
Published: (2024)
Provable Effects of Data Replay in Continual Learning: A Feature Learning Perspective
by: Ding, Meng, et al.
Published: (2026)
by: Ding, Meng, et al.
Published: (2026)
CORE: Mitigating Catastrophic Forgetting in Continual Learning through Cognitive Replay
by: Zhang, Jianshu, et al.
Published: (2024)
by: Zhang, Jianshu, et al.
Published: (2024)
Communication-Efficient Federated Learning with Accelerated Client Gradient
by: Kim, Geeho, et al.
Published: (2022)
by: Kim, Geeho, et al.
Published: (2022)
Replay4NCL: An Efficient Memory Replay-based Methodology for Neuromorphic Continual Learning in Embedded AI Systems
by: Minhas, Mishal Fatima, et al.
Published: (2025)
by: Minhas, Mishal Fatima, et al.
Published: (2025)
FedDr+: Stabilizing Dot-regression with Global Feature Distillation for Federated Learning
by: Kim, Seongyoon, et al.
Published: (2024)
by: Kim, Seongyoon, et al.
Published: (2024)
Vertex-Softmax: Tight Transformer Verification via Exact Softmax Optimization
by: Rezazadeh, Navid, et al.
Published: (2026)
by: Rezazadeh, Navid, et al.
Published: (2026)
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
by: Sheen, Heejune, et al.
Published: (2024)
by: Sheen, Heejune, et al.
Published: (2024)
Heuristic Algorithm-based Action Masking Reinforcement Learning (HAAM-RL) with Ensemble Inference Method
by: Choi, Kyuwon, et al.
Published: (2024)
by: Choi, Kyuwon, et al.
Published: (2024)
AL-GNN: Privacy-Preserving and Replay-Free Continual Graph Learning via Analytic Learning
by: Zhang, Xuling, et al.
Published: (2025)
by: Zhang, Xuling, et al.
Published: (2025)
Efficient Diversity-based Experience Replay for Deep Reinforcement Learning
by: Zhao, Kaiyan, et al.
Published: (2024)
by: Zhao, Kaiyan, et al.
Published: (2024)
Stabilizing the Q-Gradient Field for Policy Smoothness in Actor-Critic
by: Lee, Jeong Woon, et al.
Published: (2026)
by: Lee, Jeong Woon, et al.
Published: (2026)
Prioritized Trajectory Replay: A Replay Memory for Data-driven Reinforcement Learning
by: Liu, Jinyi, et al.
Published: (2023)
by: Liu, Jinyi, et al.
Published: (2023)
Scalable Strategies for Continual Learning with Replay
by: Hickok, Truman
Published: (2025)
by: Hickok, Truman
Published: (2025)
Beyond 5G Network Failure Classification for Network Digital Twin Using Graph Neural Network
by: Isah, Abubakar, et al.
Published: (2024)
by: Isah, Abubakar, et al.
Published: (2024)
CUER: Corrected Uniform Experience Replay for Off-Policy Continuous Deep Reinforcement Learning Algorithms
by: Yenicesu, Arda Sarp, et al.
Published: (2024)
by: Yenicesu, Arda Sarp, et al.
Published: (2024)
Universal Approximation with Softmax Attention
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
Revisiting Random Walks for Learning on Graphs
by: Kim, Jinwoo, et al.
Published: (2024)
by: Kim, Jinwoo, et al.
Published: (2024)
Learning Flexible Forward Trajectories for Masked Molecular Diffusion
by: Seo, Hyunjin, et al.
Published: (2025)
by: Seo, Hyunjin, et al.
Published: (2025)
To Softmax, or not to Softmax: that is the question when applying Active Learning for Transformer Models
by: Gonsior, Julius, et al.
Published: (2022)
by: Gonsior, Julius, et al.
Published: (2022)
Experience Replay Addresses Loss of Plasticity in Continual Learning
by: Wang, Jiuqi, et al.
Published: (2025)
by: Wang, Jiuqi, et al.
Published: (2025)
Minimalist Softmax Attention Provably Learns Constrained Boolean Functions
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
SplitLoRA: Balancing Stability and Plasticity in Continual Learning Through Gradient Space Splitting
by: Qiu, Haomiao, et al.
Published: (2025)
by: Qiu, Haomiao, et al.
Published: (2025)
When to Stop Reusing: Dynamic Gradient Gating for Sample-Efficient RLVR
by: Miao, Yuchun, et al.
Published: (2026)
by: Miao, Yuchun, et al.
Published: (2026)
Gradient Routing: Masking Gradients to Localize Computation in Neural Networks
by: Cloud, Alex, et al.
Published: (2024)
by: Cloud, Alex, et al.
Published: (2024)
Similar Items
-
Asymptotic Midpoint Mixup for Margin Balancing and Moderate Broadening
by: Kim, Hoyong, et al.
Published: (2024) -
Multiple Invertible and Partial-Equivariant Function for Latent Vector Transformation to Enhance Disentanglement in VAEs
by: Jung, Hee-Jun, et al.
Published: (2025) -
CFASL: Composite Factor-Aligned Symmetry Learning for Disentanglement in Variational AutoEncoder
by: Jung, Hee-Jun, et al.
Published: (2024) -
Teaching AI to Remember: Insights from Brain-Inspired Replay in Continual Learning
by: Kim, Jina
Published: (2025) -
Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models
by: Abbes, Istabrak, et al.
Published: (2025)