Efficient Continual Pre-training for Building Domain Specific Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xie, Yong, Aggarwal, Karan, Ahmad, Aitzaz |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Controlled Automatic Task-Specific Synthetic Data Generation for Hallucination Detection
von: Xie, Yong, et al.
Veröffentlicht: (2024)
von: Xie, Yong, et al.
Veröffentlicht: (2024)
Learning from Contrastive Prompts: Automated Optimization and Adaptation
von: Li, Mingqi, et al.
Veröffentlicht: (2024)
von: Li, Mingqi, et al.
Veröffentlicht: (2024)
Construction of Domain-specified Japanese Large Language Model for Finance through Continual Pre-training
von: Hirano, Masanori, et al.
Veröffentlicht: (2024)
von: Hirano, Masanori, et al.
Veröffentlicht: (2024)
Towards Effective and Efficient Continual Pre-training of Large Language Models
von: Chen, Jie, et al.
Veröffentlicht: (2024)
von: Chen, Jie, et al.
Veröffentlicht: (2024)
Examining Forgetting in Continual Pre-training of Aligned Large Language Models
von: Li, Chen-An, et al.
Veröffentlicht: (2024)
von: Li, Chen-An, et al.
Veröffentlicht: (2024)
Vi-Mistral-X: Building a Vietnamese Language Model with Advanced Continual Pre-training
von: Vo, James
Veröffentlicht: (2024)
von: Vo, James
Veröffentlicht: (2024)
Efficiently Building a Domain-Specific Large Language Model from Scratch: A Case Study of a Classical Chinese Large Language Model
von: Li, Shen, et al.
Veröffentlicht: (2025)
von: Li, Shen, et al.
Veröffentlicht: (2025)
STEP: Staged Parameter-Efficient Pre-training for Large Language Models
von: Yano, Kazuki, et al.
Veröffentlicht: (2025)
von: Yano, Kazuki, et al.
Veröffentlicht: (2025)
RASL: Retrieval Augmented Schema Linking for Massive Database Text-to-SQL
von: Eben, Jeffrey, et al.
Veröffentlicht: (2025)
von: Eben, Jeffrey, et al.
Veröffentlicht: (2025)
Efficient Continual Pre-training of LLMs for Low-resource Languages
von: Nag, Arijit, et al.
Veröffentlicht: (2024)
von: Nag, Arijit, et al.
Veröffentlicht: (2024)
Simple and Scalable Strategies to Continually Pre-train Large Language Models
von: Ibrahim, Adam, et al.
Veröffentlicht: (2024)
von: Ibrahim, Adam, et al.
Veröffentlicht: (2024)
Mecellem Models: Turkish Models Trained from Scratch and Continually Pre-trained for the Legal Domain
von: Uğur, Özgür, et al.
Veröffentlicht: (2026)
von: Uğur, Özgür, et al.
Veröffentlicht: (2026)
Efficient Continual Pre-training by Mitigating the Stability Gap
von: Guo, Yiduo, et al.
Veröffentlicht: (2024)
von: Guo, Yiduo, et al.
Veröffentlicht: (2024)
DACP: Domain-Adaptive Continual Pre-Training of Large Language Models for Phone Conversation Summarization
von: Fu, Xue-Yong, et al.
Veröffentlicht: (2025)
von: Fu, Xue-Yong, et al.
Veröffentlicht: (2025)
Efficient Knowledge Probing of Large Language Models by Adapting Pre-trained Embeddings
von: Sharma, Kartik, et al.
Veröffentlicht: (2025)
von: Sharma, Kartik, et al.
Veröffentlicht: (2025)
Can Continual Pre-training Bridge the Performance Gap between General-purpose and Specialized Language Models in the Medical Domain?
von: Doll, Niclas, et al.
Veröffentlicht: (2026)
von: Doll, Niclas, et al.
Veröffentlicht: (2026)
Model Merging in Pre-training of Large Language Models
von: Li, Yunshui, et al.
Veröffentlicht: (2025)
von: Li, Yunshui, et al.
Veröffentlicht: (2025)
Pre-trained Language Models for Keyphrase Generation: A Thorough Empirical Study
von: Wu, Di, et al.
Veröffentlicht: (2022)
von: Wu, Di, et al.
Veröffentlicht: (2022)
On Leveraging Encoder-only Pre-trained Language Models for Effective Keyphrase Generation
von: Wu, Di, et al.
Veröffentlicht: (2024)
von: Wu, Di, et al.
Veröffentlicht: (2024)
Scaling Agents via Continual Pre-training
von: Su, Liangcai, et al.
Veröffentlicht: (2025)
von: Su, Liangcai, et al.
Veröffentlicht: (2025)
Advances in Pre-trained Language Models for Domain-Specific Text Classification: A Systematic Review
von: Rostam, Zhyar Rzgar K., et al.
Veröffentlicht: (2025)
von: Rostam, Zhyar Rzgar K., et al.
Veröffentlicht: (2025)
Efficient Language Adaptive Pre-training: Extending State-of-the-Art Large Language Models for Polish
von: Ruciński, Szymon
Veröffentlicht: (2024)
von: Ruciński, Szymon
Veröffentlicht: (2024)
On the Multilingual Ability of Decoder-based Pre-trained Language Models: Finding and Controlling Language-Specific Neurons
von: Kojima, Takeshi, et al.
Veröffentlicht: (2024)
von: Kojima, Takeshi, et al.
Veröffentlicht: (2024)
D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language Models
von: Que, Haoran, et al.
Veröffentlicht: (2024)
von: Que, Haoran, et al.
Veröffentlicht: (2024)
From Curated Data to Scalable Models: Continual Pre-training of Dense and MoE Large Language Models for Tibetan
von: Yang, Lei, et al.
Veröffentlicht: (2025)
von: Yang, Lei, et al.
Veröffentlicht: (2025)
Synthesize-on-Graph: Knowledgeable Synthetic Data Generation for Continue Pre-training of Large Language Models
von: Ma, Shengjie, et al.
Veröffentlicht: (2025)
von: Ma, Shengjie, et al.
Veröffentlicht: (2025)
OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration
von: Wang, Shaobo, et al.
Veröffentlicht: (2026)
von: Wang, Shaobo, et al.
Veröffentlicht: (2026)
Exploring Continual Fine-Tuning for Enhancing Language Ability in Large Language Model
von: Aggarwal, Divyanshu, et al.
Veröffentlicht: (2024)
von: Aggarwal, Divyanshu, et al.
Veröffentlicht: (2024)
Domain-Adapted Pre-trained Language Models for Implicit Information Extraction in Crash Narratives
von: Wang, Xixi, et al.
Veröffentlicht: (2025)
von: Wang, Xixi, et al.
Veröffentlicht: (2025)
G-MAP: General Memory-Augmented Pre-trained Language Model for Domain Tasks
von: Wan, Zhongwei, et al.
Veröffentlicht: (2022)
von: Wan, Zhongwei, et al.
Veröffentlicht: (2022)
Cross-layer Attention Sharing for Pre-trained Large Language Models
von: Mu, Yongyu, et al.
Veröffentlicht: (2024)
von: Mu, Yongyu, et al.
Veröffentlicht: (2024)
Domain Pre-training Impact on Representations
von: Gonzalez-Gutierrez, Cesar, et al.
Veröffentlicht: (2025)
von: Gonzalez-Gutierrez, Cesar, et al.
Veröffentlicht: (2025)
Building Domain-Specific Small Language Models via Guided Data Generation
von: Kumar, Aman, et al.
Veröffentlicht: (2025)
von: Kumar, Aman, et al.
Veröffentlicht: (2025)
Domain-Adaptive Continued Pre-Training of Small Language Models
von: Faroz, Salman
Veröffentlicht: (2025)
von: Faroz, Salman
Veröffentlicht: (2025)
Table Comprehension in Building Codes using Vision Language Models and Domain-Specific Fine-Tuning
von: Aqib, Mohammad, et al.
Veröffentlicht: (2025)
von: Aqib, Mohammad, et al.
Veröffentlicht: (2025)
Velocitune: A Velocity-based Dynamic Domain Reweighting Method for Continual Pre-training
von: Luo, Zheheng, et al.
Veröffentlicht: (2024)
von: Luo, Zheheng, et al.
Veröffentlicht: (2024)
Towards Tracing Trustworthiness Dynamics: Revisiting Pre-training Period of Large Language Models
von: Qian, Chen, et al.
Veröffentlicht: (2024)
von: Qian, Chen, et al.
Veröffentlicht: (2024)
MIA-Tuner: Adapting Large Language Models as Pre-training Text Detector
von: Fu, Wenjie, et al.
Veröffentlicht: (2024)
von: Fu, Wenjie, et al.
Veröffentlicht: (2024)
MAPLE: Multilingual Evaluation of Parameter Efficient Finetuning of Large Language Models
von: Aggarwal, Divyanshu, et al.
Veröffentlicht: (2024)
von: Aggarwal, Divyanshu, et al.
Veröffentlicht: (2024)
Pre-training Language Model Incorporating Domain-specific Heterogeneous Knowledge into A Unified Representation
von: Zhu, Hongyin, et al.
Veröffentlicht: (2021)
von: Zhu, Hongyin, et al.
Veröffentlicht: (2021)
Ähnliche Einträge
-
Controlled Automatic Task-Specific Synthetic Data Generation for Hallucination Detection
von: Xie, Yong, et al.
Veröffentlicht: (2024) -
Learning from Contrastive Prompts: Automated Optimization and Adaptation
von: Li, Mingqi, et al.
Veröffentlicht: (2024) -
Construction of Domain-specified Japanese Large Language Model for Finance through Continual Pre-training
von: Hirano, Masanori, et al.
Veröffentlicht: (2024) -
Towards Effective and Efficient Continual Pre-training of Large Language Models
von: Chen, Jie, et al.
Veröffentlicht: (2024) -
Examining Forgetting in Continual Pre-training of Aligned Large Language Models
von: Li, Chen-An, et al.
Veröffentlicht: (2024)