Attractor Patch Networks: Reducing Catastrophic Forgetting with Routed Low-Rank Patch Experts
Fuente:
arXiv
Salvato in:
| Autore principale: | Shashank |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Analyzing and Reducing Catastrophic Forgetting in Parameter Efficient Tuning
di: Ren, Weijieying, et al.
Pubblicazione: (2024)
di: Ren, Weijieying, et al.
Pubblicazione: (2024)
Sharpness-Aware Pretraining Mitigates Catastrophic Forgetting
di: Watts, Ishaan, et al.
Pubblicazione: (2026)
di: Watts, Ishaan, et al.
Pubblicazione: (2026)
Dynamic Expert Specialization: Towards Catastrophic Forgetting-Free Multi-Domain MoE Adaptation
di: Li, Junzhuo, et al.
Pubblicazione: (2025)
di: Li, Junzhuo, et al.
Pubblicazione: (2025)
Understanding Catastrophic Forgetting in Language Models via Implicit Inference
di: Kotha, Suhas, et al.
Pubblicazione: (2023)
di: Kotha, Suhas, et al.
Pubblicazione: (2023)
Evolutionary Strategies lead to Catastrophic Forgetting in LLMs
di: Abdi, Immanuel, et al.
Pubblicazione: (2026)
di: Abdi, Immanuel, et al.
Pubblicazione: (2026)
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models
di: Zheng, Lin, et al.
Pubblicazione: (2026)
di: Zheng, Lin, et al.
Pubblicazione: (2026)
Mitigating Catastrophic Forgetting in Mathematical Reasoning Finetuning through Mixed Training
di: Reynolds, John Graham
Pubblicazione: (2025)
di: Reynolds, John Graham
Pubblicazione: (2025)
Measuring Catastrophic Forgetting in Cross-Lingual Transfer Paradigms: Exploring Tuning Strategies
di: Koloski, Boshko, et al.
Pubblicazione: (2023)
di: Koloski, Boshko, et al.
Pubblicazione: (2023)
Intelligent Learning Rate Distribution to reduce Catastrophic Forgetting in Transformers
di: Kenneweg, Philip, et al.
Pubblicazione: (2024)
di: Kenneweg, Philip, et al.
Pubblicazione: (2024)
BA-LoRA: Bias-Alleviating Low-Rank Adaptation to Mitigate Catastrophic Inheritance in Large Language Models
di: Chang, Yupeng, et al.
Pubblicazione: (2024)
di: Chang, Yupeng, et al.
Pubblicazione: (2024)
Ensembles of Low-Rank Expert Adapters
di: Li, Yinghao, et al.
Pubblicazione: (2025)
di: Li, Yinghao, et al.
Pubblicazione: (2025)
Patch Ranking: Efficient CLIP by Learning to Rank Local Patches
di: Wu, Cheng-En, et al.
Pubblicazione: (2024)
di: Wu, Cheng-En, et al.
Pubblicazione: (2024)
CURLoRA: Stable LLM Continual Fine-Tuning and Catastrophic Forgetting Mitigation
di: Fawi, Muhammad
Pubblicazione: (2024)
di: Fawi, Muhammad
Pubblicazione: (2024)
Auto-Patching: Enhancing Multi-Hop Reasoning in Language Models
di: Jan, Aviv, et al.
Pubblicazione: (2025)
di: Jan, Aviv, et al.
Pubblicazione: (2025)
Fuse to Forget: Bias Reduction and Selective Memorization through Model Fusion
di: Zaman, Kerem, et al.
Pubblicazione: (2023)
di: Zaman, Kerem, et al.
Pubblicazione: (2023)
Self-Distillation as a Performance Recovery Mechanism for LLMs: Counteracting Compression and Catastrophic Forgetting
di: Liu, Chi, et al.
Pubblicazione: (2026)
di: Liu, Chi, et al.
Pubblicazione: (2026)
DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers
di: Sharma, Aman, et al.
Pubblicazione: (2025)
di: Sharma, Aman, et al.
Pubblicazione: (2025)
Maximum Score Routing For Mixture-of-Experts
di: Dong, Bowen, et al.
Pubblicazione: (2025)
di: Dong, Bowen, et al.
Pubblicazione: (2025)
Queryable LoRA: Instruction-Regularized Routing Over Shared Low-Rank Update Atoms
di: Vaidya, Omatharv Bharat, et al.
Pubblicazione: (2026)
di: Vaidya, Omatharv Bharat, et al.
Pubblicazione: (2026)
Patch the Distribution Mismatch: RL Rewriting Agent for Stable Off-Policy SFT
di: Wang, Jiacheng, et al.
Pubblicazione: (2026)
di: Wang, Jiacheng, et al.
Pubblicazione: (2026)
RotMoLE: Enhancing Mixture of Low-Rank Experts through Rotational Gating Mechanism
di: Sun, Mengyang, et al.
Pubblicazione: (2026)
di: Sun, Mengyang, et al.
Pubblicazione: (2026)
FLoE: Fisher-Based Layer Selection for Efficient Sparse Adaptation of Low-Rank Experts
di: Wang, Xinyi, et al.
Pubblicazione: (2025)
di: Wang, Xinyi, et al.
Pubblicazione: (2025)
Measuring the Depth of LLM Unlearning via Activation Patching
di: Lee, Jaeung, et al.
Pubblicazione: (2026)
di: Lee, Jaeung, et al.
Pubblicazione: (2026)
When Are Experts Misrouted? Counterfactual Routing Analysis in Mixture-of-Experts Language Models
di: Yoon, Youngsik, et al.
Pubblicazione: (2026)
di: Yoon, Youngsik, et al.
Pubblicazione: (2026)
Dynamic Orthogonal Continual Fine-tuning for Mitigating Catastrophic Forgettings
di: Zhang, Zhixin, et al.
Pubblicazione: (2025)
di: Zhang, Zhixin, et al.
Pubblicazione: (2025)
Demystifying Language Model Forgetting with Low-rank Example Associations
di: Jin, Xisen, et al.
Pubblicazione: (2024)
di: Jin, Xisen, et al.
Pubblicazione: (2024)
Routing-Free Mixture-of-Experts
di: Liu, Yilun, et al.
Pubblicazione: (2026)
di: Liu, Yilun, et al.
Pubblicazione: (2026)
Multilingual Routing in Mixture-of-Experts
di: Bandarkar, Lucas, et al.
Pubblicazione: (2025)
di: Bandarkar, Lucas, et al.
Pubblicazione: (2025)
Dissecting Persona-Driven Reasoning in Language Models via Activation Patching
di: Poonia, Ansh, et al.
Pubblicazione: (2025)
di: Poonia, Ansh, et al.
Pubblicazione: (2025)
Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
di: Zhang, Fred, et al.
Pubblicazione: (2023)
di: Zhang, Fred, et al.
Pubblicazione: (2023)
Forecasting Time Series with LLMs via Patch-Based Prompting and Decomposition
di: Bumb, Mayank, et al.
Pubblicazione: (2025)
di: Bumb, Mayank, et al.
Pubblicazione: (2025)
CartesianMoE: Boosting Knowledge Sharing among Experts via Cartesian Product Routing in Mixture-of-Experts
di: Su, Zhenpeng, et al.
Pubblicazione: (2024)
di: Su, Zhenpeng, et al.
Pubblicazione: (2024)
Expert-Choice Routing Enables Adaptive Computation in Diffusion Language Models
di: Zhang, Shuibai, et al.
Pubblicazione: (2026)
di: Zhang, Shuibai, et al.
Pubblicazione: (2026)
On Catastrophic Forgetting in Low-Rank Decomposition-Based Parameter-Efficient Fine-Tuning
di: Ahmad, Muhammad, et al.
Pubblicazione: (2026)
di: Ahmad, Muhammad, et al.
Pubblicazione: (2026)
Patch-Prompt Aligned Bayesian Prompt Tuning for Vision-Language Models
di: Liu, Xinyang, et al.
Pubblicazione: (2023)
di: Liu, Xinyang, et al.
Pubblicazione: (2023)
Beyond Next Token Prediction: Patch-Level Training for Large Language Models
di: Shao, Chenze, et al.
Pubblicazione: (2024)
di: Shao, Chenze, et al.
Pubblicazione: (2024)
LoRI: Reducing Cross-Task Interference in Multi-Task Low-Rank Adaptation
di: Zhang, Juzheng, et al.
Pubblicazione: (2025)
di: Zhang, Juzheng, et al.
Pubblicazione: (2025)
Latent Prototype Routing: Achieving Near-Perfect Load Balancing in Mixture-of-Experts
di: Yang, Jiajie
Pubblicazione: (2025)
di: Yang, Jiajie
Pubblicazione: (2025)
Harder Tasks Need More Experts: Dynamic Routing in MoE Models
di: Huang, Quzhe, et al.
Pubblicazione: (2024)
di: Huang, Quzhe, et al.
Pubblicazione: (2024)
Sparse Memory Finetuning as a Low-Forgetting Alternative to LoRA and Full Finetuning
di: Gupta, Prakhar, et al.
Pubblicazione: (2026)
di: Gupta, Prakhar, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Analyzing and Reducing Catastrophic Forgetting in Parameter Efficient Tuning
di: Ren, Weijieying, et al.
Pubblicazione: (2024) -
Sharpness-Aware Pretraining Mitigates Catastrophic Forgetting
di: Watts, Ishaan, et al.
Pubblicazione: (2026) -
Dynamic Expert Specialization: Towards Catastrophic Forgetting-Free Multi-Domain MoE Adaptation
di: Li, Junzhuo, et al.
Pubblicazione: (2025) -
Understanding Catastrophic Forgetting in Language Models via Implicit Inference
di: Kotha, Suhas, et al.
Pubblicazione: (2023) -
Evolutionary Strategies lead to Catastrophic Forgetting in LLMs
di: Abdi, Immanuel, et al.
Pubblicazione: (2026)