Soup to go: mitigating forgetting during continual learning with model averaging
Fuente:
arXiv
Salvato in:
| Autori principali: | Kleiman, Anat, Dziugaite, Gintare Karolina, Frankle, Jonathan, Kakade, Sham, Paul, Mansheej |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Less is More: Undertraining Experts Improves Model Upcycling
di: Horoi, Stefan, et al.
Pubblicazione: (2025)
di: Horoi, Stefan, et al.
Pubblicazione: (2025)
Dataset Difficulty and the Role of Inductive Bias
di: Kwok, Devin, et al.
Pubblicazione: (2024)
di: Kwok, Devin, et al.
Pubblicazione: (2024)
Transcendence: Generative Models Can Outperform The Experts That Train Them
di: Zhang, Edwin, et al.
Pubblicazione: (2024)
di: Zhang, Edwin, et al.
Pubblicazione: (2024)
Evaluating Interventional Reasoning Capabilities of Large Language Models
di: Kasetty, Tejas, et al.
Pubblicazione: (2024)
di: Kasetty, Tejas, et al.
Pubblicazione: (2024)
Torque-Aware Momentum
di: Malviya, Pranshu, et al.
Pubblicazione: (2024)
di: Malviya, Pranshu, et al.
Pubblicazione: (2024)
SSFL: Discovering Sparse Unified Subnetworks at Initialization for Efficient Federated Learning
di: Ohib, Riyasat, et al.
Pubblicazione: (2024)
di: Ohib, Riyasat, et al.
Pubblicazione: (2024)
Continual Learning in Vision-Language Models via Aligned Model Merging
di: Sokar, Ghada, et al.
Pubblicazione: (2025)
di: Sokar, Ghada, et al.
Pubblicazione: (2025)
Connections between Schedule-Free Optimizers, AdEMAMix, and Accelerated SGD Variants
di: Morwani, Depen, et al.
Pubblicazione: (2025)
di: Morwani, Depen, et al.
Pubblicazione: (2025)
The Potential of Second-Order Optimization for LLMs: A Study with Full Gauss-Newton
di: Abreu, Natalie, et al.
Pubblicazione: (2025)
di: Abreu, Natalie, et al.
Pubblicazione: (2025)
From Dormant to Deleted: Tamper-Resistant Unlearning Through Weight-Space Regularization
di: Siddiqui, Shoaib Ahmed, et al.
Pubblicazione: (2025)
di: Siddiqui, Shoaib Ahmed, et al.
Pubblicazione: (2025)
Mixtures of Experts Unlock Parameter Scaling for Deep RL
di: Obando-Ceron, Johan, et al.
Pubblicazione: (2024)
di: Obando-Ceron, Johan, et al.
Pubblicazione: (2024)
Learning Hidden Markov Models Using Conditional Samples
di: Kakade, Sham M., et al.
Pubblicazione: (2023)
di: Kakade, Sham M., et al.
Pubblicazione: (2023)
Flash Inference: Near Linear Time Inference for Long Convolution Sequence Models and Beyond
di: Oncescu, Costin-Andrei, et al.
Pubblicazione: (2024)
di: Oncescu, Costin-Andrei, et al.
Pubblicazione: (2024)
Deconstructing What Makes a Good Optimizer for Language Models
di: Zhao, Rosie, et al.
Pubblicazione: (2024)
di: Zhao, Rosie, et al.
Pubblicazione: (2024)
CoLoR-Filter: Conditional Loss Reduction Filtering for Targeted Language Model Pre-training
di: Brandfonbrener, David, et al.
Pubblicazione: (2024)
di: Brandfonbrener, David, et al.
Pubblicazione: (2024)
Task diversity produces systematic transfer but inhibits continual reinforcement learning
di: Seth, Purab, et al.
Pubblicazione: (2026)
di: Seth, Purab, et al.
Pubblicazione: (2026)
Discovering Hierarchical Latent Capabilities of Language Models via Causal Representation Learning
di: Jin, Jikai, et al.
Pubblicazione: (2025)
di: Jin, Jikai, et al.
Pubblicazione: (2025)
Prescriptive Scaling Reveals the Evolution of Language Model Capabilities
di: Zhang, Hanlin, et al.
Pubblicazione: (2026)
di: Zhang, Hanlin, et al.
Pubblicazione: (2026)
The Non-Local Model Merging Problem: Permutation Symmetries and Variance Collapse
di: Sharma, Ekansh, et al.
Pubblicazione: (2024)
di: Sharma, Ekansh, et al.
Pubblicazione: (2024)
Repeat After Me: Transformers are Better than State Space Models at Copying
di: Jelassi, Samy, et al.
Pubblicazione: (2024)
di: Jelassi, Samy, et al.
Pubblicazione: (2024)
Mixture of Experts in a Mixture of RL settings
di: Willi, Timon, et al.
Pubblicazione: (2024)
di: Willi, Timon, et al.
Pubblicazione: (2024)
Out-of-distribution forgetting: vulnerability of continual learning to intra-class distribution shift
di: Guo, Liangxuan, et al.
Pubblicazione: (2023)
di: Guo, Liangxuan, et al.
Pubblicazione: (2023)
Scaling Laws for Imitation Learning in Single-Agent Games
di: Tuyls, Jens, et al.
Pubblicazione: (2023)
di: Tuyls, Jens, et al.
Pubblicazione: (2023)
Loss-to-Loss Prediction: Scaling Laws for All Datasets
di: Brandfonbrener, David, et al.
Pubblicazione: (2024)
di: Brandfonbrener, David, et al.
Pubblicazione: (2024)
LoRA Learns Less and Forgets Less
di: Biderman, Dan, et al.
Pubblicazione: (2024)
di: Biderman, Dan, et al.
Pubblicazione: (2024)
The Role of Sparsity for Length Generalization in Transformers
di: Golowich, Noah, et al.
Pubblicazione: (2025)
di: Golowich, Noah, et al.
Pubblicazione: (2025)
Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
di: Meterez, Alexandru, et al.
Pubblicazione: (2026)
di: Meterez, Alexandru, et al.
Pubblicazione: (2026)
Does your data spark joy? Performance gains from domain upsampling at the end of training
di: Blakeney, Cody, et al.
Pubblicazione: (2024)
di: Blakeney, Cody, et al.
Pubblicazione: (2024)
Data Selection for Transfer Unlearning
di: Sepahvand, Nazanin Mohammadi, et al.
Pubblicazione: (2024)
di: Sepahvand, Nazanin Mohammadi, et al.
Pubblicazione: (2024)
Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems
di: Qi, Zhenting, et al.
Pubblicazione: (2024)
di: Qi, Zhenting, et al.
Pubblicazione: (2024)
SOAP: Improving and Stabilizing Shampoo using Adam
di: Vyas, Nikhil, et al.
Pubblicazione: (2024)
di: Vyas, Nikhil, et al.
Pubblicazione: (2024)
State Soup: In-Context Skill Learning, Retrieval and Mixing
di: Pióro, Maciej, et al.
Pubblicazione: (2024)
di: Pióro, Maciej, et al.
Pubblicazione: (2024)
Identifying Spurious Biases Early in Training through the Lens of Simplicity Bias
di: Yang, Yu, et al.
Pubblicazione: (2023)
di: Yang, Yu, et al.
Pubblicazione: (2023)
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
di: Meterez, Alexandru, et al.
Pubblicazione: (2025)
di: Meterez, Alexandru, et al.
Pubblicazione: (2025)
Scaling Laws in Linear Regression: Compute, Parameters, and Data
di: Lin, Licong, et al.
Pubblicazione: (2024)
di: Lin, Licong, et al.
Pubblicazione: (2024)
Bone Soups: A Seek-and-Soup Model Merging Approach for Controllable Multi-Objective Generation
di: Xie, Guofu, et al.
Pubblicazione: (2025)
di: Xie, Guofu, et al.
Pubblicazione: (2025)
The SMeL Test: A simple benchmark for media literacy in language models
di: Ahdritz, Gustaf, et al.
Pubblicazione: (2025)
di: Ahdritz, Gustaf, et al.
Pubblicazione: (2025)
Leveraging Function Space Aggregation for Federated Learning at Scale
di: Dhawan, Nikita, et al.
Pubblicazione: (2023)
di: Dhawan, Nikita, et al.
Pubblicazione: (2023)
Unlearning in- vs. out-of-distribution data in LLMs under gradient-based method
di: Baluta, Teodora, et al.
Pubblicazione: (2024)
di: Baluta, Teodora, et al.
Pubblicazione: (2024)
ECG-Soup: Harnessing Multi-Layer Synergy for ECG Foundation Models
di: Nguyen, Phu X., et al.
Pubblicazione: (2025)
di: Nguyen, Phu X., et al.
Pubblicazione: (2025)
Documenti analoghi
-
Less is More: Undertraining Experts Improves Model Upcycling
di: Horoi, Stefan, et al.
Pubblicazione: (2025) -
Dataset Difficulty and the Role of Inductive Bias
di: Kwok, Devin, et al.
Pubblicazione: (2024) -
Transcendence: Generative Models Can Outperform The Experts That Train Them
di: Zhang, Edwin, et al.
Pubblicazione: (2024) -
Evaluating Interventional Reasoning Capabilities of Large Language Models
di: Kasetty, Tejas, et al.
Pubblicazione: (2024) -
Torque-Aware Momentum
di: Malviya, Pranshu, et al.
Pubblicazione: (2024)