Fine-Tune an SLM or Prompt an LLM? The Case of Generating Low-Code Workflows

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ayala, Orlando Marquez, Bechard, Patrice, Chen, Emily, Baird, Maggie, Chen, Jingfei
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915394571206656
author Ayala, Orlando Marquez
Bechard, Patrice
Chen, Emily
Baird, Maggie
Chen, Jingfei
author_facet Ayala, Orlando Marquez
Bechard, Patrice
Chen, Emily
Baird, Maggie
Chen, Jingfei
contents Large Language Models (LLMs) such as GPT-4o can handle a wide range of complex tasks with the right prompt. As per token costs are reduced, the advantages of fine-tuning Small Language Models (SLMs) for real-world applications -- faster inference, lower costs -- may no longer be clear. In this work, we present evidence that, for domain-specific tasks that require structured outputs, SLMs still have a quality advantage. We compare fine-tuning an SLM against prompting LLMs on the task of generating low-code workflows in JSON form. We observe that while a good prompt can yield reasonable results, fine-tuning improves quality by 10% on average. We also perform systematic error analysis to reveal model limitations.
format Preprint
id arxiv_https___arxiv_org_abs_2505_24189
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fine-Tune an SLM or Prompt an LLM? The Case of Generating Low-Code Workflows
Ayala, Orlando Marquez
Bechard, Patrice
Chen, Emily
Baird, Maggie
Chen, Jingfei
Machine Learning
Artificial Intelligence
Computation and Language
Large Language Models (LLMs) such as GPT-4o can handle a wide range of complex tasks with the right prompt. As per token costs are reduced, the advantages of fine-tuning Small Language Models (SLMs) for real-world applications -- faster inference, lower costs -- may no longer be clear. In this work, we present evidence that, for domain-specific tasks that require structured outputs, SLMs still have a quality advantage. We compare fine-tuning an SLM against prompting LLMs on the task of generating low-code workflows in JSON form. We observe that while a good prompt can yield reasonable results, fine-tuning improves quality by 10% on average. We also perform systematic error analysis to reveal model limitations.
title Fine-Tune an SLM or Prompt an LLM? The Case of Generating Low-Code Workflows
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2505.24189