Dual-Phase Continual Learning: Supervised Adaptation Meets Unsupervised Retention
Fuente:
arXiv
Guardado en:
| Autores principales: | Singh, Vaibhav, Aljundi, Rahaf, Belilovsky, Eugene |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
When Data Falls Short: Grokking Below the Critical Threshold
por: Singh, Vaibhav, et al.
Publicado: (2025)
por: Singh, Vaibhav, et al.
Publicado: (2025)
Beyond Cosine Decay: On the effectiveness of Infinite Learning Rate Schedule for Continual Pre-training
por: Singh, Vaibhav, et al.
Publicado: (2025)
por: Singh, Vaibhav, et al.
Publicado: (2025)
Model Parallelism With Subnetwork Data Parallelism
por: Singh, Vaibhav, et al.
Publicado: (2025)
por: Singh, Vaibhav, et al.
Publicado: (2025)
DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone
por: Singh, Vaibhav, et al.
Publicado: (2025)
por: Singh, Vaibhav, et al.
Publicado: (2025)
Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL
por: Miahi, Erfan, et al.
Publicado: (2026)
por: Miahi, Erfan, et al.
Publicado: (2026)
FairDropout: Using Example-Tied Dropout to Enhance Generalization of Minority Groups
por: Nanfack, Geraldin, et al.
Publicado: (2025)
por: Nanfack, Geraldin, et al.
Publicado: (2025)
Recurrent Attention-based Token Selection for Efficient Streaming Video-LLMs
por: Dorovatas, Vaggelis, et al.
Publicado: (2025)
por: Dorovatas, Vaggelis, et al.
Publicado: (2025)
Model Breadcrumbs: Scaling Multi-Task Model Merging with Sparse Masks
por: Davari, MohammadReza, et al.
Publicado: (2023)
por: Davari, MohammadReza, et al.
Publicado: (2023)
Celo2: Towards Learned Optimization Free Lunch
por: Moudgil, Abhinav, et al.
Publicado: (2026)
por: Moudgil, Abhinav, et al.
Publicado: (2026)
Stabilizing Native Low-Rank LLM Pretraining
por: Janson, Paul, et al.
Publicado: (2026)
por: Janson, Paul, et al.
Publicado: (2026)
Celo: Training Versatile Learned Optimizers on a Compute Diet
por: Moudgil, Abhinav, et al.
Publicado: (2025)
por: Moudgil, Abhinav, et al.
Publicado: (2025)
Test Time Adaptation Using Adaptive Quantile Recalibration
por: Mehrbod, Paria, et al.
Publicado: (2025)
por: Mehrbod, Paria, et al.
Publicado: (2025)
Warming Up for Zeroth-Order Federated Pre-Training with Low Resource Clients
por: Legate, Gwen, et al.
Publicado: (2025)
por: Legate, Gwen, et al.
Publicado: (2025)
Efficient Refusal Ablation in LLM through Optimal Transport
por: Nanfack, Geraldin, et al.
Publicado: (2026)
por: Nanfack, Geraldin, et al.
Publicado: (2026)
Communication Efficient LLM Pre-training with SparseLoCo
por: Sarfi, Amir, et al.
Publicado: (2025)
por: Sarfi, Amir, et al.
Publicado: (2025)
Not Only the Last-Layer Features for Spurious Correlations: All Layer Deep Feature Reweighting
por: Hameed, Humza Wajid, et al.
Publicado: (2024)
por: Hameed, Humza Wajid, et al.
Publicado: (2024)
Meta-learning Optimizers for Communication-Efficient Learning
por: Joseph, Charles-Étienne, et al.
Publicado: (2023)
por: Joseph, Charles-Étienne, et al.
Publicado: (2023)
Imperfect Vision Encoders: Efficient and Robust Tuning for Vision-Language Models
por: Panos, Aristeidis, et al.
Publicado: (2024)
por: Panos, Aristeidis, et al.
Publicado: (2024)
DragD3D: Realistic Mesh Editing with Rigidity Control Driven by 2D Diffusion Priors
por: Xie, Tianhao, et al.
Publicado: (2023)
por: Xie, Tianhao, et al.
Publicado: (2023)
Rethinking Prompt Optimization: Reinforcement, Diversification, and Migration in Blackbox LLMs
por: Davari, MohammadReza, et al.
Publicado: (2025)
por: Davari, MohammadReza, et al.
Publicado: (2025)
From Feature Visualization to Visual Circuits: Effect of Adversarial Model Manipulation
por: Nanfack, Geraldin, et al.
Publicado: (2024)
por: Nanfack, Geraldin, et al.
Publicado: (2024)
PyLO: Towards Accessible Learned Optimizers in PyTorch
por: Janson, Paul, et al.
Publicado: (2025)
por: Janson, Paul, et al.
Publicado: (2025)
SEM: Sparse Embedding Modulation for Post-Hoc Debiasing of Vision-Language Models
por: Guimard, Quentin, et al.
Publicado: (2026)
por: Guimard, Quentin, et al.
Publicado: (2026)
$μ$LO: Compute-Efficient Meta-Generalization of Learned Optimizers
por: Thérien, Benjamin, et al.
Publicado: (2024)
por: Thérien, Benjamin, et al.
Publicado: (2024)
Heterogeneous Low-Bandwidth Pre-Training of LLMs
por: Obeidi, Yazan, et al.
Publicado: (2026)
por: Obeidi, Yazan, et al.
Publicado: (2026)
MuLoCo: Muon is a practical inner optimizer for DiLoCo
por: Thérien, Benjamin, et al.
Publicado: (2025)
por: Thérien, Benjamin, et al.
Publicado: (2025)
Less is More: Undertraining Experts Improves Model Upcycling
por: Horoi, Stefan, et al.
Publicado: (2025)
por: Horoi, Stefan, et al.
Publicado: (2025)
Non-Uniform Parameter-Wise Model Merging
por: Camacho, Albert Manuel Orozco, et al.
Publicado: (2024)
por: Camacho, Albert Manuel Orozco, et al.
Publicado: (2024)
Incentivizing Permissionless Distributed Learning of LLMs
por: Lidin, Joel, et al.
Publicado: (2025)
por: Lidin, Joel, et al.
Publicado: (2025)
Is Supervised Learning Really That Different from Unsupervised?
por: Allerbo, Oskar, et al.
Publicado: (2025)
por: Allerbo, Oskar, et al.
Publicado: (2025)
PETRA: Parallel End-to-end Training with Reversible Architectures
por: Rivaud, Stéphane, et al.
Publicado: (2024)
por: Rivaud, Stéphane, et al.
Publicado: (2024)
Towards Unsupervised Open-Set Graph Domain Adaptation via Dual Reprogramming
por: Zhang, Zhen, et al.
Publicado: (2025)
por: Zhang, Zhen, et al.
Publicado: (2025)
When Test-Time Adaptation Meets Self-Supervised Models
por: Han, Jisu, et al.
Publicado: (2025)
por: Han, Jisu, et al.
Publicado: (2025)
AdaFisher: Adaptive Second Order Optimization via Fisher Information
por: Gomes, Damien Martins, et al.
Publicado: (2024)
por: Gomes, Damien Martins, et al.
Publicado: (2024)
Accelerating Training with Neuron Interaction and Nowcasting Networks
por: Knyazev, Boris, et al.
Publicado: (2024)
por: Knyazev, Boris, et al.
Publicado: (2024)
Active Continual Learning: On Balancing Knowledge Retention and Learnability
por: Vu, Thuy-Trang, et al.
Publicado: (2023)
por: Vu, Thuy-Trang, et al.
Publicado: (2023)
Knowledge Retention for Continual Model-Based Reinforcement Learning
por: Sun, Yixiang, et al.
Publicado: (2025)
por: Sun, Yixiang, et al.
Publicado: (2025)
Simple and Scalable Strategies to Continually Pre-train Large Language Models
por: Ibrahim, Adam, et al.
Publicado: (2024)
por: Ibrahim, Adam, et al.
Publicado: (2024)
Quantum Unsupervised and Supervised Learning on Superconducting Processors
por: Sarma, Abhijat, et al.
Publicado: (2019)
por: Sarma, Abhijat, et al.
Publicado: (2019)
ACCO: Accumulate While You Communicate for Communication-Overlapped Sharded LLM Training
por: Nabli, Adel, et al.
Publicado: (2024)
por: Nabli, Adel, et al.
Publicado: (2024)
Ejemplares similares
-
When Data Falls Short: Grokking Below the Critical Threshold
por: Singh, Vaibhav, et al.
Publicado: (2025) -
Beyond Cosine Decay: On the effectiveness of Infinite Learning Rate Schedule for Continual Pre-training
por: Singh, Vaibhav, et al.
Publicado: (2025) -
Model Parallelism With Subnetwork Data Parallelism
por: Singh, Vaibhav, et al.
Publicado: (2025) -
DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone
por: Singh, Vaibhav, et al.
Publicado: (2025) -
Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL
por: Miahi, Erfan, et al.
Publicado: (2026)