Barriers for Learning in an Evolving World: Mathematical Understanding of Loss of Plasticity
Fuente:
arXiv
Saved in:
| Main Authors: | Joudaki, Amir, Lanzillotta, Giulia, Razlighi, Mohammad Samragh, Mirzadeh, Iman, Alizadeh, Keivan, Hofmann, Thomas, Farajtabar, Mehrdad, Faghri, Fartash |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization
by: Samragh, Mohammad, et al.
Published: (2024)
by: Samragh, Mohammad, et al.
Published: (2024)
Computational Bottlenecks of Training Small-scale Large Language Models
by: Ashkboos, Saleh, et al.
Published: (2024)
by: Ashkboos, Saleh, et al.
Published: (2024)
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
by: Mirzadeh, Iman, et al.
Published: (2024)
by: Mirzadeh, Iman, et al.
Published: (2024)
The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
by: Shojaee, Parshin, et al.
Published: (2025)
by: Shojaee, Parshin, et al.
Published: (2025)
SALSA: Soup-based Alignment Learning for Stronger Adaptation in RLHF
by: Chegini, Atoosa, et al.
Published: (2024)
by: Chegini, Atoosa, et al.
Published: (2024)
LLM in a flash: Efficient Large Language Model Inference with Limited Memory
by: Alizadeh, Keivan, et al.
Published: (2023)
by: Alizadeh, Keivan, et al.
Published: (2023)
Duo-LLM: A Framework for Studying Adaptive Computation in Large Language Models
by: Alizadeh, Keivan, et al.
Published: (2024)
by: Alizadeh, Keivan, et al.
Published: (2024)
TiC-LM: A Web-Scale Benchmark for Time-Continual LLM Pretraining
by: Li, Jeffrey, et al.
Published: (2025)
by: Li, Jeffrey, et al.
Published: (2025)
Knowledge Transfer from Vision Foundation Models for Efficient Training of Small Task-specific Models
by: Vemulapalli, Raviteja, et al.
Published: (2023)
by: Vemulapalli, Raviteja, et al.
Published: (2023)
Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context
by: Alizadeh, Keivan, et al.
Published: (2026)
by: Alizadeh, Keivan, et al.
Published: (2026)
MoE-PHDS: One MoE checkpoint for flexible runtime sparsity
by: Hannah, Lauren. A, et al.
Published: (2025)
by: Hannah, Lauren. A, et al.
Published: (2025)
CatLIP: CLIP-level Visual Recognition Accuracy with 2.7x Faster Pre-training on Web-scale Image-Text Data
by: Mehta, Sachin, et al.
Published: (2024)
by: Mehta, Sachin, et al.
Published: (2024)
SAM-CLIP: Merging Vision Foundation Models towards Semantic and Spatial Understanding
by: Wang, Haoxiang, et al.
Published: (2023)
by: Wang, Haoxiang, et al.
Published: (2023)
TiC-CLIP: Continual Training of CLIP Models
by: Garg, Saurabh, et al.
Published: (2023)
by: Garg, Saurabh, et al.
Published: (2023)
Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential
by: Samragh, Mohammad, et al.
Published: (2025)
by: Samragh, Mohammad, et al.
Published: (2025)
Proxy-FDA: Proxy-based Feature Distribution Alignment for Fine-tuning Vision Foundation Models without Forgetting
by: Huang, Chen, et al.
Published: (2025)
by: Huang, Chen, et al.
Published: (2025)
Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why
by: Armandpour, Mohammadreza, et al.
Published: (2026)
by: Armandpour, Mohammadreza, et al.
Published: (2026)
Heads collapse, features stay: Why Replay needs big buffers
by: Lanzillotta, Giulia, et al.
Published: (2025)
by: Lanzillotta, Giulia, et al.
Published: (2025)
Emergence of Globally Attracting Fixed Points in Deep Neural Networks With Nonlinear Activations
by: Joudaki, Amir, et al.
Published: (2024)
by: Joudaki, Amir, et al.
Published: (2024)
The hybrid dilemma -- do hybrid technologies play a transitionary or stationary role in transitions processes?
by: Phirouzabadi, Amir Mirzadeh
Published: (2025)
by: Phirouzabadi, Amir Mirzadeh
Published: (2025)
The Importance of Being Lazy: Scaling Limits of Continual Learning
by: Graldi, Jacopo, et al.
Published: (2025)
by: Graldi, Jacopo, et al.
Published: (2025)
CLIP with Quality Captions: A Strong Pretraining for Vision Tasks
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2024)
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2024)
MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2023)
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2023)
Local vs Global continual learning
by: Lanzillotta, Giulia, et al.
Published: (2024)
by: Lanzillotta, Giulia, et al.
Published: (2024)
Revisiting Knowledge Distillation: The Hidden Role of Dataset Size
by: Lanzillotta, Giulia, et al.
Published: (2025)
by: Lanzillotta, Giulia, et al.
Published: (2025)
Algunas consideraciones heurísticas en torno a la formación de grupos intelectuales en el Territorio Nacional de la Pampa
by: María Lanzillotta
Published: (2010)
by: María Lanzillotta
Published: (2010)
The path towards contact-based physical human-robot interaction
by: Farajtabar, Mohammad, et al.
Published: (2024)
by: Farajtabar, Mohammad, et al.
Published: (2024)
AMUSE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding
by: Chowdhury, Sanjoy, et al.
Published: (2025)
by: Chowdhury, Sanjoy, et al.
Published: (2025)
SpecMD: A Comprehensive Study On Speculative Expert Prefetching
by: Hoang, Duc, et al.
Published: (2026)
by: Hoang, Duc, et al.
Published: (2026)
SpecHop: Continuous Speculation for Accelerating Multi-Hop Retrieval Agents
by: Saberi, Mehrdad, et al.
Published: (2026)
by: Saberi, Mehrdad, et al.
Published: (2026)
MobileCLIP2: Improving Multi-Modal Reinforced Training
by: Faghri, Fartash, et al.
Published: (2025)
by: Faghri, Fartash, et al.
Published: (2025)
MUSCLE: A Model Update Strategy for Compatible LLM Evolution
by: Echterhoff, Jessica, et al.
Published: (2024)
by: Echterhoff, Jessica, et al.
Published: (2024)
Mathematical Approach in Hybrid Beamforming for ISAC Systems
by: Khosroshahi, Keivan, et al.
Published: (2025)
by: Khosroshahi, Keivan, et al.
Published: (2025)
Test Time Training for AC Power Flow Surrogates via Physics and Operational Constraint Refinement
by: Dogoulis, Panteleimon, et al.
Published: (2025)
by: Dogoulis, Panteleimon, et al.
Published: (2025)
Achieving Fine‐Grained Microstructure in Low‐Alloy Steel: A Study on Static Recrystallization Using Experimental and Simulation Approaches
by: Mahdiyeh Baharvand, et al.
Published: (2025)
by: Mahdiyeh Baharvand, et al.
Published: (2025)
Confident Splatting: Confidence-Based Compression of 3D Gaussian Splatting via Learnable Beta Distributions
by: Razlighi, AmirHossein Naghi, et al.
Published: (2025)
by: Razlighi, AmirHossein Naghi, et al.
Published: (2025)
Towards the net zero carbon future: A review of blockchain‐enabled peer‐to‐peer carbon trading
by: Mohammad Parhamfar, et al.
Published: (2024)
by: Mohammad Parhamfar, et al.
Published: (2024)
Harnessing the Ecological and Genomic Adaptability of the Bacterial Genus Massilia for Environmental and Industrial Applications
by: Kamyar Amirhosseini, et al.
Published: (2025)
by: Kamyar Amirhosseini, et al.
Published: (2025)
Reactivation: Empirical NTK Dynamics Under Task Shifts
by: Liu, Yuzhi, et al.
Published: (2025)
by: Liu, Yuzhi, et al.
Published: (2025)
Improving vision-inspired keyword spotting using dynamic module skipping in streaming conformer encoder
by: Bittar, Alexandre, et al.
Published: (2023)
by: Bittar, Alexandre, et al.
Published: (2023)
Similar Items
-
Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization
by: Samragh, Mohammad, et al.
Published: (2024) -
Computational Bottlenecks of Training Small-scale Large Language Models
by: Ashkboos, Saleh, et al.
Published: (2024) -
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
by: Mirzadeh, Iman, et al.
Published: (2024) -
The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
by: Shojaee, Parshin, et al.
Published: (2025) -
SALSA: Soup-based Alignment Learning for Stronger Adaptation in RLHF
by: Chegini, Atoosa, et al.
Published: (2024)