Fine-Tune an SLM or Prompt an LLM? The Case of Generating Low-Code Workflows
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866915394571206656 |
|---|---|
| author | Ayala, Orlando Marquez Bechard, Patrice Chen, Emily Baird, Maggie Chen, Jingfei |
| author_facet | Ayala, Orlando Marquez Bechard, Patrice Chen, Emily Baird, Maggie Chen, Jingfei |
| contents | Large Language Models (LLMs) such as GPT-4o can handle a wide range of complex tasks with the right prompt. As per token costs are reduced, the advantages of fine-tuning Small Language Models (SLMs) for real-world applications -- faster inference, lower costs -- may no longer be clear. In this work, we present evidence that, for domain-specific tasks that require structured outputs, SLMs still have a quality advantage. We compare fine-tuning an SLM against prompting LLMs on the task of generating low-code workflows in JSON form. We observe that while a good prompt can yield reasonable results, fine-tuning improves quality by 10% on average. We also perform systematic error analysis to reveal model limitations. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_24189 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Fine-Tune an SLM or Prompt an LLM? The Case of Generating Low-Code Workflows Ayala, Orlando Marquez Bechard, Patrice Chen, Emily Baird, Maggie Chen, Jingfei Machine Learning Artificial Intelligence Computation and Language Large Language Models (LLMs) such as GPT-4o can handle a wide range of complex tasks with the right prompt. As per token costs are reduced, the advantages of fine-tuning Small Language Models (SLMs) for real-world applications -- faster inference, lower costs -- may no longer be clear. In this work, we present evidence that, for domain-specific tasks that require structured outputs, SLMs still have a quality advantage. We compare fine-tuning an SLM against prompting LLMs on the task of generating low-code workflows in JSON form. We observe that while a good prompt can yield reasonable results, fine-tuning improves quality by 10% on average. We also perform systematic error analysis to reveal model limitations. |
| title | Fine-Tune an SLM or Prompt an LLM? The Case of Generating Low-Code Workflows |
| topic | Machine Learning Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2505.24189 |