Intelligent Learning Rate Distribution to reduce Catastrophic Forgetting in Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Kenneweg, Philip, Schulz, Alexander, Schröder, Sarah, Hammer, Barbara |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Faster Convergence for Transformer Fine-tuning with Line Search Methods
by: Kenneweg, Philip, et al.
Published: (2024)
by: Kenneweg, Philip, et al.
Published: (2024)
Improving Line Search Methods for Large Scale Neural Network Training
by: Kenneweg, Philip, et al.
Published: (2024)
by: Kenneweg, Philip, et al.
Published: (2024)
No learning rates needed: Introducing SALSA -- Stable Armijo Line Search Adaptation
by: Kenneweg, Philip, et al.
Published: (2024)
by: Kenneweg, Philip, et al.
Published: (2024)
Evolutionary Strategies lead to Catastrophic Forgetting in LLMs
by: Abdi, Immanuel, et al.
Published: (2026)
by: Abdi, Immanuel, et al.
Published: (2026)
Debiasing Sentence Embedders through Contrastive Word Pairs
by: Kenneweg, Philip, et al.
Published: (2024)
by: Kenneweg, Philip, et al.
Published: (2024)
Targeted Visualization of the Backbone of Encoder LLMs
by: Roberts, Isaac, et al.
Published: (2024)
by: Roberts, Isaac, et al.
Published: (2024)
Analyzing and Reducing Catastrophic Forgetting in Parameter Efficient Tuning
by: Ren, Weijieying, et al.
Published: (2024)
by: Ren, Weijieying, et al.
Published: (2024)
Neural Architecture Search for Sentence Classification with BERT
by: Kenneweg, Philip, et al.
Published: (2024)
by: Kenneweg, Philip, et al.
Published: (2024)
CURLoRA: Stable LLM Continual Fine-Tuning and Catastrophic Forgetting Mitigation
by: Fawi, Muhammad
Published: (2024)
by: Fawi, Muhammad
Published: (2024)
Self-Distillation as a Performance Recovery Mechanism for LLMs: Counteracting Compression and Catastrophic Forgetting
by: Liu, Chi, et al.
Published: (2026)
by: Liu, Chi, et al.
Published: (2026)
Dynamic Expert Specialization: Towards Catastrophic Forgetting-Free Multi-Domain MoE Adaptation
by: Li, Junzhuo, et al.
Published: (2025)
by: Li, Junzhuo, et al.
Published: (2025)
Forgetting Transformer: Softmax Attention with a Forget Gate
by: Lin, Zhixuan, et al.
Published: (2025)
by: Lin, Zhixuan, et al.
Published: (2025)
Adaptive Computation Pruning for the Forgetting Transformer
by: Lin, Zhixuan, et al.
Published: (2025)
by: Lin, Zhixuan, et al.
Published: (2025)
LoRA Learns Less and Forgets Less
by: Biderman, Dan, et al.
Published: (2024)
by: Biderman, Dan, et al.
Published: (2024)
Dynamic Orthogonal Continual Fine-tuning for Mitigating Catastrophic Forgettings
by: Zhang, Zhixin, et al.
Published: (2025)
by: Zhang, Zhixin, et al.
Published: (2025)
Evaluating Metrics for Bias in Word Embeddings
by: Schröder, Sarah, et al.
Published: (2021)
by: Schröder, Sarah, et al.
Published: (2021)
FOREVER: Forgetting Curve-Inspired Memory Replay for Language Model Continual Learning
by: Feng, Yujie, et al.
Published: (2026)
by: Feng, Yujie, et al.
Published: (2026)
Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models
by: Hsiao, Chi-Yuan, et al.
Published: (2025)
by: Hsiao, Chi-Yuan, et al.
Published: (2025)
Offline Learning and Forgetting for Reasoning with Large Language Models
by: Ni, Tianwei, et al.
Published: (2025)
by: Ni, Tianwei, et al.
Published: (2025)
Continual Learning and Catastrophic Forgetting
by: van de Ven, Gido M., et al.
Published: (2024)
by: van de Ven, Gido M., et al.
Published: (2024)
Retrieval Augmented Generation Systems: Automatic Dataset Creation, Evaluation and Boolean Agent Setup
by: Kenneweg, Tristan, et al.
Published: (2024)
by: Kenneweg, Tristan, et al.
Published: (2024)
Wings: Learning Multimodal LLMs without Text-only Forgetting
by: Zhang, Yi-Kai, et al.
Published: (2024)
by: Zhang, Yi-Kai, et al.
Published: (2024)
Real Time Detection and Quantitative Analysis of Spurious Forgetting in Continual Learning
by: Wang, Weiwei
Published: (2025)
by: Wang, Weiwei
Published: (2025)
Transformers Learn to Achieve Second-Order Convergence Rates for In-Context Linear Regression
by: Fu, Deqing, et al.
Published: (2023)
by: Fu, Deqing, et al.
Published: (2023)
Elephants Never Forget: Memorization and Learning of Tabular Data in Large Language Models
by: Bordt, Sebastian, et al.
Published: (2024)
by: Bordt, Sebastian, et al.
Published: (2024)
Sequencing to Mitigate Catastrophic Forgetting in Continual Learning
by: Moussa, Hesham G., et al.
Published: (2025)
by: Moussa, Hesham G., et al.
Published: (2025)
Self-Training for Sample-Efficient Active Learning for Text Classification with Pre-Trained Language Models
by: Schröder, Christopher, et al.
Published: (2024)
by: Schröder, Christopher, et al.
Published: (2024)
Graceful Forgetting in Generative Language Models
by: Jiang, Chunyang, et al.
Published: (2025)
by: Jiang, Chunyang, et al.
Published: (2025)
Gradient Correlation Subspace Learning against Catastrophic Forgetting
by: Dubnov, Tammuz, et al.
Published: (2024)
by: Dubnov, Tammuz, et al.
Published: (2024)
Negative Preference Optimization: From Catastrophic Collapse to Effective Unlearning
by: Zhang, Ruiqi, et al.
Published: (2024)
by: Zhang, Ruiqi, et al.
Published: (2024)
Learn More, Forget Less: A Gradient-Aware Data Selection Approach for LLM
by: Liu, Yibai, et al.
Published: (2025)
by: Liu, Yibai, et al.
Published: (2025)
Stable Preference Optimization: A Bilevel Approach to Catastrophic Preference Shift
by: Jian, Chengtao, et al.
Published: (2025)
by: Jian, Chengtao, et al.
Published: (2025)
Reference-Free Rating of LLM Responses via Latent Information
by: Girrbach, Leander, et al.
Published: (2025)
by: Girrbach, Leander, et al.
Published: (2025)
Unveiling and Addressing Pseudo Forgetting in Large Language Models
by: Sun, Huashan, et al.
Published: (2024)
by: Sun, Huashan, et al.
Published: (2024)
Stuffed Mamba: Oversized States Lead to the Inability to Forget
by: Chen, Yingfa, et al.
Published: (2024)
by: Chen, Yingfa, et al.
Published: (2024)
Mapping Post-Training Forgetting in Language Models at Scale
by: Harmon, Jackson, et al.
Published: (2025)
by: Harmon, Jackson, et al.
Published: (2025)
Addressing Loss of Plasticity and Catastrophic Forgetting in Continual Learning
by: Elsayed, Mohamed, et al.
Published: (2024)
by: Elsayed, Mohamed, et al.
Published: (2024)
On the Implicit Adversariality of Catastrophic Forgetting in Deep Continual Learning
by: Peng, Ze, et al.
Published: (2025)
by: Peng, Ze, et al.
Published: (2025)
Learning is Forgetting: LLM Training As Lossy Compression
by: Conklin, Henry C., et al.
Published: (2026)
by: Conklin, Henry C., et al.
Published: (2026)
GeRe: Towards Efficient Anti-Forgetting in Continual Learning of LLM via General Samples Replay
by: Zhang, Yunan, et al.
Published: (2025)
by: Zhang, Yunan, et al.
Published: (2025)
Similar Items
-
Faster Convergence for Transformer Fine-tuning with Line Search Methods
by: Kenneweg, Philip, et al.
Published: (2024) -
Improving Line Search Methods for Large Scale Neural Network Training
by: Kenneweg, Philip, et al.
Published: (2024) -
No learning rates needed: Introducing SALSA -- Stable Armijo Line Search Adaptation
by: Kenneweg, Philip, et al.
Published: (2024) -
Evolutionary Strategies lead to Catastrophic Forgetting in LLMs
by: Abdi, Immanuel, et al.
Published: (2026) -
Debiasing Sentence Embedders through Contrastive Word Pairs
by: Kenneweg, Philip, et al.
Published: (2024)