NLD-LLM: A systematic framework for evaluating small language transformer models on natural language description

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jelodar, Hamed, Meymani, Mohammad, Hamedi, Parisa, Nwankwo, Tochukwu Emmanuel, Bai, Samita, Razavi-Far, Roozbeh, Ghorbani, Ali A.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908578890121216
author Jelodar, Hamed
Meymani, Mohammad
Hamedi, Parisa
Nwankwo, Tochukwu Emmanuel
Bai, Samita
Razavi-Far, Roozbeh
Ghorbani, Ali A.
author_facet Jelodar, Hamed
Meymani, Mohammad
Hamedi, Parisa
Nwankwo, Tochukwu Emmanuel
Bai, Samita
Razavi-Far, Roozbeh
Ghorbani, Ali A.
contents Natural Language Description (NLD) is a Natural Language Processing (NLP) task that requires models to generate structured and meaningful outputs from natural language inputs. In this work, we propose NLD-LLM, a systematic NLP framework to evaluate the performance of language models to generate accurate and concise source code descriptions. This framework incorporates a diverse set of transformer models, including Qwen, DeepSeek, Phi, LLaMA, and Mistral, spanning various sizes, architectures, and training approaches. Central to NLD-LLM is a comprehensive prompt design strategy that includes standardized formatting, clear task guidance, and NLD prompting, ensuring fair and consistent evaluation. Additionally, we apply an iterative refinement process to improve output's quality and assess the model's adaptability. Using semantic and structural metrics, our analysis demonstrates that prompt engineering significantly impacts the effectiveness of the model such that smaller models often performing competitively when supported by well-crafted prompts.
format Preprint
id arxiv_https___arxiv_org_abs_2510_05139
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NLD-LLM: A systematic framework for evaluating small language transformer models on natural language description
Jelodar, Hamed
Meymani, Mohammad
Hamedi, Parisa
Nwankwo, Tochukwu Emmanuel
Bai, Samita
Razavi-Far, Roozbeh
Ghorbani, Ali A.
Computation and Language
Natural Language Description (NLD) is a Natural Language Processing (NLP) task that requires models to generate structured and meaningful outputs from natural language inputs. In this work, we propose NLD-LLM, a systematic NLP framework to evaluate the performance of language models to generate accurate and concise source code descriptions. This framework incorporates a diverse set of transformer models, including Qwen, DeepSeek, Phi, LLaMA, and Mistral, spanning various sizes, architectures, and training approaches. Central to NLD-LLM is a comprehensive prompt design strategy that includes standardized formatting, clear task guidance, and NLD prompting, ensuring fair and consistent evaluation. Additionally, we apply an iterative refinement process to improve output's quality and assess the model's adaptability. Using semantic and structural metrics, our analysis demonstrates that prompt engineering significantly impacts the effectiveness of the model such that smaller models often performing competitively when supported by well-crafted prompts.
title NLD-LLM: A systematic framework for evaluating small language transformer models on natural language description
topic Computation and Language
url https://arxiv.org/abs/2510.05139