Building Domain-Specific Small Language Models via Guided Data Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kumar, Aman, Amin, Ekant Muljibhai, Lee, Xian Yeow, Vidyaratne, Lasitha, Farahat, Ahmed K., Ghosh, Dipanjan D., Koreeda, Yuta, Gupta, Chetan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909927953399808
author Kumar, Aman
Amin, Ekant Muljibhai
Lee, Xian Yeow
Vidyaratne, Lasitha
Farahat, Ahmed K.
Ghosh, Dipanjan D.
Koreeda, Yuta
Gupta, Chetan
author_facet Kumar, Aman
Amin, Ekant Muljibhai
Lee, Xian Yeow
Vidyaratne, Lasitha
Farahat, Ahmed K.
Ghosh, Dipanjan D.
Koreeda, Yuta
Gupta, Chetan
contents Large Language Models (LLMs) have shown remarkable success in supporting a wide range of knowledge-intensive tasks. In specialized domains, there is growing interest in leveraging LLMs to assist subject matter experts with domain-specific challenges. However, deploying LLMs as SaaS solutions raises data privacy concerns, while many open-source models demand significant computational resources for effective domain adaptation and deployment. A promising alternative is to develop smaller, domain-specialized LLMs, though this approach is often constrained by the lack of high-quality domain-specific training data. In this work, we address these limitations by presenting a cost-efficient and scalable training pipeline that combines guided synthetic data generation from a small seed corpus with bottom-up domain data curation. Our pipeline integrates Domain-Adaptive Pretraining (DAPT), Domain-specific Supervised Fine-tuning (DSFT), and Direct Preference Optimization (DPO) to train effective small-scale models for specialized use cases. We demonstrate this approach through DiagnosticSLM, a 3B-parameter domain-specific model tailored for fault diagnosis, root cause analysis, and repair recommendation in industrial settings. To evaluate model performance, we introduce four domain-specific benchmarks: multiple-choice questions (DiagnosticMCQ), question answering (DiagnosticQA), sentence completion (DiagnosticComp), and summarization (DiagnosticSum). DiagnosticSLM achieves up to 25% accuracy improvement over open-source models of comparable or larger size (2B-9B) on the MCQ task, while also outperforming or matching them in other tasks, demonstrating effective domain-specific reasoning and generalization capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2511_21748
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Building Domain-Specific Small Language Models via Guided Data Generation
Kumar, Aman
Amin, Ekant Muljibhai
Lee, Xian Yeow
Vidyaratne, Lasitha
Farahat, Ahmed K.
Ghosh, Dipanjan D.
Koreeda, Yuta
Gupta, Chetan
Computation and Language
Artificial Intelligence
Large Language Models (LLMs) have shown remarkable success in supporting a wide range of knowledge-intensive tasks. In specialized domains, there is growing interest in leveraging LLMs to assist subject matter experts with domain-specific challenges. However, deploying LLMs as SaaS solutions raises data privacy concerns, while many open-source models demand significant computational resources for effective domain adaptation and deployment. A promising alternative is to develop smaller, domain-specialized LLMs, though this approach is often constrained by the lack of high-quality domain-specific training data. In this work, we address these limitations by presenting a cost-efficient and scalable training pipeline that combines guided synthetic data generation from a small seed corpus with bottom-up domain data curation. Our pipeline integrates Domain-Adaptive Pretraining (DAPT), Domain-specific Supervised Fine-tuning (DSFT), and Direct Preference Optimization (DPO) to train effective small-scale models for specialized use cases. We demonstrate this approach through DiagnosticSLM, a 3B-parameter domain-specific model tailored for fault diagnosis, root cause analysis, and repair recommendation in industrial settings. To evaluate model performance, we introduce four domain-specific benchmarks: multiple-choice questions (DiagnosticMCQ), question answering (DiagnosticQA), sentence completion (DiagnosticComp), and summarization (DiagnosticSum). DiagnosticSLM achieves up to 25% accuracy improvement over open-source models of comparable or larger size (2B-9B) on the MCQ task, while also outperforming or matching them in other tasks, demonstrating effective domain-specific reasoning and generalization capabilities.
title Building Domain-Specific Small Language Models via Guided Data Generation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2511.21748