Forgetting Transformer: Softmax Attention with a Forget Gate
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Zhixuan, Nikishin, Evgenii, He, Xu Owen, Courville, Aaron |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Adaptive Computation Pruning for the Forgetting Transformer
von: Lin, Zhixuan, et al.
Veröffentlicht: (2025)
von: Lin, Zhixuan, et al.
Veröffentlicht: (2025)
Intelligent Learning Rate Distribution to reduce Catastrophic Forgetting in Transformers
von: Kenneweg, Philip, et al.
Veröffentlicht: (2024)
von: Kenneweg, Philip, et al.
Veröffentlicht: (2024)
In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
von: Collins, Liam, et al.
Veröffentlicht: (2024)
von: Collins, Liam, et al.
Veröffentlicht: (2024)
Graceful Forgetting in Generative Language Models
von: Jiang, Chunyang, et al.
Veröffentlicht: (2025)
von: Jiang, Chunyang, et al.
Veröffentlicht: (2025)
Stuffed Mamba: Oversized States Lead to the Inability to Forget
von: Chen, Yingfa, et al.
Veröffentlicht: (2024)
von: Chen, Yingfa, et al.
Veröffentlicht: (2024)
Scalable-Softmax Is Superior for Attention
von: Nakanishi, Ken M.
Veröffentlicht: (2025)
von: Nakanishi, Ken M.
Veröffentlicht: (2025)
The Curse of Diversity in Ensemble-Based Exploration
von: Lin, Zhixuan, et al.
Veröffentlicht: (2024)
von: Lin, Zhixuan, et al.
Veröffentlicht: (2024)
Evolutionary Strategies lead to Catastrophic Forgetting in LLMs
von: Abdi, Immanuel, et al.
Veröffentlicht: (2026)
von: Abdi, Immanuel, et al.
Veröffentlicht: (2026)
LoRA Learns Less and Forgets Less
von: Biderman, Dan, et al.
Veröffentlicht: (2024)
von: Biderman, Dan, et al.
Veröffentlicht: (2024)
Scaling Stick-Breaking Attention: An Efficient Implementation and In-depth Study
von: Tan, Shawn, et al.
Veröffentlicht: (2024)
von: Tan, Shawn, et al.
Veröffentlicht: (2024)
Wings: Learning Multimodal LLMs without Text-only Forgetting
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2024)
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2024)
Offline Learning and Forgetting for Reasoning with Large Language Models
von: Ni, Tianwei, et al.
Veröffentlicht: (2025)
von: Ni, Tianwei, et al.
Veröffentlicht: (2025)
Mapping Post-Training Forgetting in Language Models at Scale
von: Harmon, Jackson, et al.
Veröffentlicht: (2025)
von: Harmon, Jackson, et al.
Veröffentlicht: (2025)
Unveiling and Addressing Pseudo Forgetting in Large Language Models
von: Sun, Huashan, et al.
Veröffentlicht: (2024)
von: Sun, Huashan, et al.
Veröffentlicht: (2024)
Analyzing and Reducing Catastrophic Forgetting in Parameter Efficient Tuning
von: Ren, Weijieying, et al.
Veröffentlicht: (2024)
von: Ren, Weijieying, et al.
Veröffentlicht: (2024)
Self-Distillation as a Performance Recovery Mechanism for LLMs: Counteracting Compression and Catastrophic Forgetting
von: Liu, Chi, et al.
Veröffentlicht: (2026)
von: Liu, Chi, et al.
Veröffentlicht: (2026)
Entropy-Adaptive Fine-Tuning: Resolving Confident Conflicts to Mitigate Forgetting
von: Diao, Muxi, et al.
Veröffentlicht: (2026)
von: Diao, Muxi, et al.
Veröffentlicht: (2026)
Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training
von: Lai, Song, et al.
Veröffentlicht: (2025)
von: Lai, Song, et al.
Veröffentlicht: (2025)
How Much Can We Forget about Data Contamination?
von: Bordt, Sebastian, et al.
Veröffentlicht: (2024)
von: Bordt, Sebastian, et al.
Veröffentlicht: (2024)
LLM Unlearning via Loss Adjustment with Only Forget Data
von: Wang, Yaxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yaxuan, et al.
Veröffentlicht: (2024)
FOREVER: Forgetting Curve-Inspired Memory Replay for Language Model Continual Learning
von: Feng, Yujie, et al.
Veröffentlicht: (2026)
von: Feng, Yujie, et al.
Veröffentlicht: (2026)
To Softmax, or not to Softmax: that is the question when applying Active Learning for Transformer Models
von: Gonsior, Julius, et al.
Veröffentlicht: (2022)
von: Gonsior, Julius, et al.
Veröffentlicht: (2022)
Don't Forget Imagination!
von: Vityaev, Evgenii E., et al.
Veröffentlicht: (2025)
von: Vityaev, Evgenii E., et al.
Veröffentlicht: (2025)
Real Time Detection and Quantitative Analysis of Spurious Forgetting in Continual Learning
von: Wang, Weiwei
Veröffentlicht: (2025)
von: Wang, Weiwei
Veröffentlicht: (2025)
Forget What Matters, Keep the Rest: Selective Unlearning of Informative Tokens
von: Koh, Seunghee, et al.
Veröffentlicht: (2026)
von: Koh, Seunghee, et al.
Veröffentlicht: (2026)
Fuse to Forget: Bias Reduction and Selective Memorization through Model Fusion
von: Zaman, Kerem, et al.
Veröffentlicht: (2023)
von: Zaman, Kerem, et al.
Veröffentlicht: (2023)
Forget Attention: Importance-Aware Attention Is All You Need
von: Shin, Soohyeong, et al.
Veröffentlicht: (2026)
von: Shin, Soohyeong, et al.
Veröffentlicht: (2026)
Forget What You Know about LLMs Evaluations -- LLMs are Like a Chameleon
von: Cohen-Inger, Nurit, et al.
Veröffentlicht: (2025)
von: Cohen-Inger, Nurit, et al.
Veröffentlicht: (2025)
Learning is Forgetting: LLM Training As Lossy Compression
von: Conklin, Henry C., et al.
Veröffentlicht: (2026)
von: Conklin, Henry C., et al.
Veröffentlicht: (2026)
CURLoRA: Stable LLM Continual Fine-Tuning and Catastrophic Forgetting Mitigation
von: Fawi, Muhammad
Veröffentlicht: (2024)
von: Fawi, Muhammad
Veröffentlicht: (2024)
Elephants Never Forget: Memorization and Learning of Tabular Data in Large Language Models
von: Bordt, Sebastian, et al.
Veröffentlicht: (2024)
von: Bordt, Sebastian, et al.
Veröffentlicht: (2024)
FIT to Forget: Robust Continual Unlearning for Large Language Models
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2026)
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2026)
Learn More, Forget Less: A Gradient-Aware Data Selection Approach for LLM
von: Liu, Yibai, et al.
Veröffentlicht: (2025)
von: Liu, Yibai, et al.
Veröffentlicht: (2025)
BAPO: Base-Anchored Preference Optimization for Overcoming Forgetting in Large Language Models Personalization
von: Lee, Gihun, et al.
Veröffentlicht: (2024)
von: Lee, Gihun, et al.
Veröffentlicht: (2024)
Don't Forget to Connect! Improving RAG with Graph-based Reranking
von: Dong, Jialin, et al.
Veröffentlicht: (2024)
von: Dong, Jialin, et al.
Veröffentlicht: (2024)
Dynamic Expert Specialization: Towards Catastrophic Forgetting-Free Multi-Domain MoE Adaptation
von: Li, Junzhuo, et al.
Veröffentlicht: (2025)
von: Li, Junzhuo, et al.
Veröffentlicht: (2025)
Did You Forget What I Asked? Prospective Memory Failures in Large Language Models
von: Mittal, Avni
Veröffentlicht: (2026)
von: Mittal, Avni
Veröffentlicht: (2026)
GeRe: Towards Efficient Anti-Forgetting in Continual Learning of LLM via General Samples Replay
von: Zhang, Yunan, et al.
Veröffentlicht: (2025)
von: Zhang, Yunan, et al.
Veröffentlicht: (2025)
Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning
von: Shang, Bingqi, et al.
Veröffentlicht: (2025)
von: Shang, Bingqi, et al.
Veröffentlicht: (2025)
Forgetting-MarI: LLM Unlearning via Marginal Information Regularization
von: Xu, Shizhou, et al.
Veröffentlicht: (2025)
von: Xu, Shizhou, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Adaptive Computation Pruning for the Forgetting Transformer
von: Lin, Zhixuan, et al.
Veröffentlicht: (2025) -
Intelligent Learning Rate Distribution to reduce Catastrophic Forgetting in Transformers
von: Kenneweg, Philip, et al.
Veröffentlicht: (2024) -
In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
von: Collins, Liam, et al.
Veröffentlicht: (2024) -
Graceful Forgetting in Generative Language Models
von: Jiang, Chunyang, et al.
Veröffentlicht: (2025) -
Stuffed Mamba: Oversized States Lead to the Inability to Forget
von: Chen, Yingfa, et al.
Veröffentlicht: (2024)