Rehearsal-Free Modular and Compositional Continual Learning for Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Mingyang, Adel, Heike, Lange, Lukas, Strötgen, Jannik, Schütze, Hinrich |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learn it or Leave it: Module Composition and Pruning for Continual Learning
von: Wang, Mingyang, et al.
Veröffentlicht: (2024)
von: Wang, Mingyang, et al.
Veröffentlicht: (2024)
Better Call SAUL: Fluent and Consistent Language Model Editing with Generation Regularization
von: Wang, Mingyang, et al.
Veröffentlicht: (2024)
von: Wang, Mingyang, et al.
Veröffentlicht: (2024)
NLNDE at SemEval-2023 Task 12: Adaptive Pretraining and Source Language Selection for Low-Resource Multilingual Sentiment Analysis
von: Wang, Mingyang, et al.
Veröffentlicht: (2023)
von: Wang, Mingyang, et al.
Veröffentlicht: (2023)
Language Mixing in Reasoning Language Models: Patterns, Impact, and Internal Causes
von: Wang, Mingyang, et al.
Veröffentlicht: (2025)
von: Wang, Mingyang, et al.
Veröffentlicht: (2025)
Bring Your Own Knowledge: A Survey of Methods for LLM Knowledge Expansion
von: Wang, Mingyang, et al.
Veröffentlicht: (2025)
von: Wang, Mingyang, et al.
Veröffentlicht: (2025)
Lost in Multilinguality: Dissecting Cross-lingual Factual Inconsistency in Transformer Language Models
von: Wang, Mingyang, et al.
Veröffentlicht: (2025)
von: Wang, Mingyang, et al.
Veröffentlicht: (2025)
Discourse-Aware In-Context Learning for Temporal Expression Normalization
von: Gautam, Akash Kumar, et al.
Veröffentlicht: (2024)
von: Gautam, Akash Kumar, et al.
Veröffentlicht: (2024)
GLUScope: A Tool for Analyzing GLU Neurons in Transformer Language Models
von: Gerstner, Sebastian, et al.
Veröffentlicht: (2026)
von: Gerstner, Sebastian, et al.
Veröffentlicht: (2026)
Understanding Gated Neurons in Transformers from Their Input-Output Functionality
von: Gerstner, Sebastian, et al.
Veröffentlicht: (2025)
von: Gerstner, Sebastian, et al.
Veröffentlicht: (2025)
DATA: Decomposed Attention-based Task Adaptation for Rehearsal-Free Continual Learning
von: Liao, Huanxuan, et al.
Veröffentlicht: (2025)
von: Liao, Huanxuan, et al.
Veröffentlicht: (2025)
Your Pretrained Model Tells the Difficulty Itself: A Self-Adaptive Curriculum Learning Paradigm for Natural Language Understanding
von: Feng, Qi, et al.
Veröffentlicht: (2025)
von: Feng, Qi, et al.
Veröffentlicht: (2025)
HYPEROFA: Expanding LLM Vocabulary to New Languages via Hypernetwork-Based Embedding Initialization
von: Özeren, Enes, et al.
Veröffentlicht: (2025)
von: Özeren, Enes, et al.
Veröffentlicht: (2025)
Through a Compressed Lens: Investigating The Impact of Quantization on Factual Knowledge Recall
von: Wang, Qianli, et al.
Veröffentlicht: (2025)
von: Wang, Qianli, et al.
Veröffentlicht: (2025)
Derivational Morphology Reveals Analogical Generalization in Large Language Models
von: Hofmann, Valentin, et al.
Veröffentlicht: (2024)
von: Hofmann, Valentin, et al.
Veröffentlicht: (2024)
GRASP: A Rehearsal Policy for Efficient Online Continual Learning
von: Harun, Md Yousuf, et al.
Veröffentlicht: (2023)
von: Harun, Md Yousuf, et al.
Veröffentlicht: (2023)
ChunkFT: Byte-Streamed Optimization for Memory-Efficient Full Fine-Tuning
von: Liu, Yongkang, et al.
Veröffentlicht: (2026)
von: Liu, Yongkang, et al.
Veröffentlicht: (2026)
BlackboxNLP-2025 MIB Shared Task: Exploring Ensemble Strategies for Circuit Localization Methods
von: Mondorf, Philipp, et al.
Veröffentlicht: (2025)
von: Mondorf, Philipp, et al.
Veröffentlicht: (2025)
CRAFT Your Dataset: Task-Specific Synthetic Dataset Generation Through Corpus Retrieval and Augmentation
von: Ziegler, Ingo, et al.
Veröffentlicht: (2024)
von: Ziegler, Ingo, et al.
Veröffentlicht: (2024)
LongForm: Effective Instruction Tuning with Reverse Instructions
von: Köksal, Abdullatif, et al.
Veröffentlicht: (2023)
von: Köksal, Abdullatif, et al.
Veröffentlicht: (2023)
OFA: A Framework of Initializing Unseen Subword Embeddings for Efficient Large-scale Multilingual Continued Pretraining
von: Liu, Yihong, et al.
Veröffentlicht: (2023)
von: Liu, Yihong, et al.
Veröffentlicht: (2023)
The Anatomy of an Edit: Mechanism-Guided Activation Steering for Knowledge Editing
von: Cao, Yuan, et al.
Veröffentlicht: (2026)
von: Cao, Yuan, et al.
Veröffentlicht: (2026)
MURI: High-Quality Instruction Tuning Datasets for Low-Resource Languages via Reverse Instructions
von: Köksal, Abdullatif, et al.
Veröffentlicht: (2024)
von: Köksal, Abdullatif, et al.
Veröffentlicht: (2024)
Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
PromptDSI: Prompt-based Rehearsal-free Continual Learning for Document Retrieval
von: Huynh, Tuan-Luc, et al.
Veröffentlicht: (2024)
von: Huynh, Tuan-Luc, et al.
Veröffentlicht: (2024)
Consistent Prompting for Rehearsal-Free Continual Learning
von: Gao, Zhanxin, et al.
Veröffentlicht: (2024)
von: Gao, Zhanxin, et al.
Veröffentlicht: (2024)
Rehearsal-Free Continual Federated Learning with Synergistic Synaptic Intelligence
von: Li, Yichen, et al.
Veröffentlicht: (2024)
von: Li, Yichen, et al.
Veröffentlicht: (2024)
A Survey of On-Policy Distillation for Large Language Models
von: Song, Mingyang, et al.
Veröffentlicht: (2026)
von: Song, Mingyang, et al.
Veröffentlicht: (2026)
Reflecting on the State of Rehearsal-free Continual Learning with Pretrained Models
von: Thede, Lukas, et al.
Veröffentlicht: (2024)
von: Thede, Lukas, et al.
Veröffentlicht: (2024)
HiFT: A Hierarchical Full Parameter Fine-Tuning Strategy
von: Liu, Yongkang, et al.
Veröffentlicht: (2024)
von: Liu, Yongkang, et al.
Veröffentlicht: (2024)
Selective Self-Rehearsal: A Fine-Tuning Approach to Improve Generalization in Large Language Models
von: Gupta, Sonam, et al.
Veröffentlicht: (2024)
von: Gupta, Sonam, et al.
Veröffentlicht: (2024)
Learning to Route for Dynamic Adapter Composition in Continual Learning with Language Models
von: Araujo, Vladimir, et al.
Veröffentlicht: (2024)
von: Araujo, Vladimir, et al.
Veröffentlicht: (2024)
Thought Flow Nets: From Single Predictions to Trains of Model Thought
von: Schuff, Hendrik, et al.
Veröffentlicht: (2021)
von: Schuff, Hendrik, et al.
Veröffentlicht: (2021)
Refusal Direction is Universal Across Safety-Aligned Languages
von: Wang, Xinpeng, et al.
Veröffentlicht: (2025)
von: Wang, Xinpeng, et al.
Veröffentlicht: (2025)
Steering MoE LLMs via Expert (De)Activation
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2025)
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2025)
LangSAMP: Language-Script Aware Multilingual Pretraining
von: Liu, Yihong, et al.
Veröffentlicht: (2024)
von: Liu, Yihong, et al.
Veröffentlicht: (2024)
SMoA: Spectrum Modulation Adapter for Parameter-Efficient Fine-Tuning
von: Liu, Yongkang, et al.
Veröffentlicht: (2026)
von: Liu, Yongkang, et al.
Veröffentlicht: (2026)
On the Entity-Level Alignment in Crosslingual Consistency
von: Liu, Yihong, et al.
Veröffentlicht: (2025)
von: Liu, Yihong, et al.
Veröffentlicht: (2025)
Fine-Tuning Large Language Models to Appropriately Abstain with Semantic Entropy
von: Tjandra, Benedict Aaron, et al.
Veröffentlicht: (2024)
von: Tjandra, Benedict Aaron, et al.
Veröffentlicht: (2024)
Composition of Experts: A Modular Compound AI System Leveraging Large Language Models
von: Jain, Swayambhoo, et al.
Veröffentlicht: (2024)
von: Jain, Swayambhoo, et al.
Veröffentlicht: (2024)
Analyzing German Parliamentary Speeches: A Machine Learning Approach for Topic and Sentiment Classification
von: Pätz, Lukas, et al.
Veröffentlicht: (2025)
von: Pätz, Lukas, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Learn it or Leave it: Module Composition and Pruning for Continual Learning
von: Wang, Mingyang, et al.
Veröffentlicht: (2024) -
Better Call SAUL: Fluent and Consistent Language Model Editing with Generation Regularization
von: Wang, Mingyang, et al.
Veröffentlicht: (2024) -
NLNDE at SemEval-2023 Task 12: Adaptive Pretraining and Source Language Selection for Low-Resource Multilingual Sentiment Analysis
von: Wang, Mingyang, et al.
Veröffentlicht: (2023) -
Language Mixing in Reasoning Language Models: Patterns, Impact, and Internal Causes
von: Wang, Mingyang, et al.
Veröffentlicht: (2025) -
Bring Your Own Knowledge: A Survey of Methods for LLM Knowledge Expansion
von: Wang, Mingyang, et al.
Veröffentlicht: (2025)