Learn it or Leave it: Module Composition and Pruning for Continual Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Mingyang, Adel, Heike, Lange, Lukas, Strötgen, Jannik, Schütze, Hinrich |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rehearsal-Free Modular and Compositional Continual Learning for Language Models
by: Wang, Mingyang, et al.
Published: (2024)
by: Wang, Mingyang, et al.
Published: (2024)
Better Call SAUL: Fluent and Consistent Language Model Editing with Generation Regularization
by: Wang, Mingyang, et al.
Published: (2024)
by: Wang, Mingyang, et al.
Published: (2024)
NLNDE at SemEval-2023 Task 12: Adaptive Pretraining and Source Language Selection for Low-Resource Multilingual Sentiment Analysis
by: Wang, Mingyang, et al.
Published: (2023)
by: Wang, Mingyang, et al.
Published: (2023)
Language Mixing in Reasoning Language Models: Patterns, Impact, and Internal Causes
by: Wang, Mingyang, et al.
Published: (2025)
by: Wang, Mingyang, et al.
Published: (2025)
Bring Your Own Knowledge: A Survey of Methods for LLM Knowledge Expansion
by: Wang, Mingyang, et al.
Published: (2025)
by: Wang, Mingyang, et al.
Published: (2025)
Lost in Multilinguality: Dissecting Cross-lingual Factual Inconsistency in Transformer Language Models
by: Wang, Mingyang, et al.
Published: (2025)
by: Wang, Mingyang, et al.
Published: (2025)
Discourse-Aware In-Context Learning for Temporal Expression Normalization
by: Gautam, Akash Kumar, et al.
Published: (2024)
by: Gautam, Akash Kumar, et al.
Published: (2024)
GLUScope: A Tool for Analyzing GLU Neurons in Transformer Language Models
by: Gerstner, Sebastian, et al.
Published: (2026)
by: Gerstner, Sebastian, et al.
Published: (2026)
Understanding Gated Neurons in Transformers from Their Input-Output Functionality
by: Gerstner, Sebastian, et al.
Published: (2025)
by: Gerstner, Sebastian, et al.
Published: (2025)
Through a Compressed Lens: Investigating The Impact of Quantization on Factual Knowledge Recall
by: Wang, Qianli, et al.
Published: (2025)
by: Wang, Qianli, et al.
Published: (2025)
Your Pretrained Model Tells the Difficulty Itself: A Self-Adaptive Curriculum Learning Paradigm for Natural Language Understanding
by: Feng, Qi, et al.
Published: (2025)
by: Feng, Qi, et al.
Published: (2025)
HYPEROFA: Expanding LLM Vocabulary to New Languages via Hypernetwork-Based Embedding Initialization
by: Özeren, Enes, et al.
Published: (2025)
by: Özeren, Enes, et al.
Published: (2025)
SMoA: Spectrum Modulation Adapter for Parameter-Efficient Fine-Tuning
by: Liu, Yongkang, et al.
Published: (2026)
by: Liu, Yongkang, et al.
Published: (2026)
ChunkFT: Byte-Streamed Optimization for Memory-Efficient Full Fine-Tuning
by: Liu, Yongkang, et al.
Published: (2026)
by: Liu, Yongkang, et al.
Published: (2026)
BlackboxNLP-2025 MIB Shared Task: Exploring Ensemble Strategies for Circuit Localization Methods
by: Mondorf, Philipp, et al.
Published: (2025)
by: Mondorf, Philipp, et al.
Published: (2025)
CRAFT Your Dataset: Task-Specific Synthetic Dataset Generation Through Corpus Retrieval and Augmentation
by: Ziegler, Ingo, et al.
Published: (2024)
by: Ziegler, Ingo, et al.
Published: (2024)
LongForm: Effective Instruction Tuning with Reverse Instructions
by: Köksal, Abdullatif, et al.
Published: (2023)
by: Köksal, Abdullatif, et al.
Published: (2023)
OFA: A Framework of Initializing Unseen Subword Embeddings for Efficient Large-scale Multilingual Continued Pretraining
by: Liu, Yihong, et al.
Published: (2023)
by: Liu, Yihong, et al.
Published: (2023)
The Anatomy of an Edit: Mechanism-Guided Activation Steering for Knowledge Editing
by: Cao, Yuan, et al.
Published: (2026)
by: Cao, Yuan, et al.
Published: (2026)
In-Context Learning Learns Label Relationships but Is Not Conventional Learning
by: Kossen, Jannik, et al.
Published: (2023)
by: Kossen, Jannik, et al.
Published: (2023)
Derivational Morphology Reveals Analogical Generalization in Large Language Models
by: Hofmann, Valentin, et al.
Published: (2024)
by: Hofmann, Valentin, et al.
Published: (2024)
Analyzing German Parliamentary Speeches: A Machine Learning Approach for Topic and Sentiment Classification
by: Pätz, Lukas, et al.
Published: (2025)
by: Pätz, Lukas, et al.
Published: (2025)
HiFT: A Hierarchical Full Parameter Fine-Tuning Strategy
by: Liu, Yongkang, et al.
Published: (2024)
by: Liu, Yongkang, et al.
Published: (2024)
MURI: High-Quality Instruction Tuning Datasets for Low-Resource Languages via Reverse Instructions
by: Köksal, Abdullatif, et al.
Published: (2024)
by: Köksal, Abdullatif, et al.
Published: (2024)
Steering MoE LLMs via Expert (De)Activation
by: Fayyaz, Mohsen, et al.
Published: (2025)
by: Fayyaz, Mohsen, et al.
Published: (2025)
On the Entity-Level Alignment in Crosslingual Consistency
by: Liu, Yihong, et al.
Published: (2025)
by: Liu, Yihong, et al.
Published: (2025)
Cut Your Losses! Learning to Prune Paths Early for Efficient Parallel Reasoning
by: Bi, Jiaxi, et al.
Published: (2026)
by: Bi, Jiaxi, et al.
Published: (2026)
Mitigating Copy Bias in In-Context Learning through Neuron Pruning
by: Ali, Ameen, et al.
Published: (2024)
by: Ali, Ameen, et al.
Published: (2024)
Learning to Route for Dynamic Adapter Composition in Continual Learning with Language Models
by: Araujo, Vladimir, et al.
Published: (2024)
by: Araujo, Vladimir, et al.
Published: (2024)
Language Model-Driven Data Pruning Enables Efficient Active Learning
by: Azeemi, Abdul Hameed, et al.
Published: (2024)
by: Azeemi, Abdul Hameed, et al.
Published: (2024)
Thought Flow Nets: From Single Predictions to Trains of Model Thought
by: Schuff, Hendrik, et al.
Published: (2021)
by: Schuff, Hendrik, et al.
Published: (2021)
Learning to Learn for Few-shot Continual Active Learning
by: Ho, Stella, et al.
Published: (2023)
by: Ho, Stella, et al.
Published: (2023)
Towards Compositionality in Concept Learning
by: Stein, Adam, et al.
Published: (2024)
by: Stein, Adam, et al.
Published: (2024)
TalkTag: Fine-Grained Morphosyntactic Error Annotation for Transcribed Speech
by: Venturini, Shamira, et al.
Published: (2026)
by: Venturini, Shamira, et al.
Published: (2026)
BMIKE-53: Investigating Cross-Lingual Knowledge Editing with In-Context Learning
by: Nie, Ercong, et al.
Published: (2024)
by: Nie, Ercong, et al.
Published: (2024)
Refusal Direction is Universal Across Safety-Aligned Languages
by: Wang, Xinpeng, et al.
Published: (2025)
by: Wang, Xinpeng, et al.
Published: (2025)
COPAL: Continual Pruning in Large Language Generative Models
by: Malla, Srikanth, et al.
Published: (2024)
by: Malla, Srikanth, et al.
Published: (2024)
STUN: Structured-Then-Unstructured Pruning for Scalable MoE Pruning
by: Lee, Jaeseong, et al.
Published: (2024)
by: Lee, Jaeseong, et al.
Published: (2024)
Pruning Weights but Not Truth: Safeguarding Truthfulness While Pruning LLMs
by: Fu, Yao, et al.
Published: (2025)
by: Fu, Yao, et al.
Published: (2025)
Continual Learning Under Language Shift
by: Gogoulou, Evangelia, et al.
Published: (2023)
by: Gogoulou, Evangelia, et al.
Published: (2023)
Similar Items
-
Rehearsal-Free Modular and Compositional Continual Learning for Language Models
by: Wang, Mingyang, et al.
Published: (2024) -
Better Call SAUL: Fluent and Consistent Language Model Editing with Generation Regularization
by: Wang, Mingyang, et al.
Published: (2024) -
NLNDE at SemEval-2023 Task 12: Adaptive Pretraining and Source Language Selection for Low-Resource Multilingual Sentiment Analysis
by: Wang, Mingyang, et al.
Published: (2023) -
Language Mixing in Reasoning Language Models: Patterns, Impact, and Internal Causes
by: Wang, Mingyang, et al.
Published: (2025) -
Bring Your Own Knowledge: A Survey of Methods for LLM Knowledge Expansion
by: Wang, Mingyang, et al.
Published: (2025)