A Comprehensive Study on the Use of Word Embedding Models in Software Engineering Domain

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Xiaohan, Zou, Weiqin, Zhi, Lianyi, Meng, Qianshuang, Zhang, Jingxuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910964317683712
author Chen, Xiaohan
Zou, Weiqin
Zhi, Lianyi
Meng, Qianshuang
Zhang, Jingxuan
author_facet Chen, Xiaohan
Zou, Weiqin
Zhi, Lianyi
Meng, Qianshuang
Zhang, Jingxuan
contents Word embedding (WE) techniques are advanced textual semantic representation models oriented from the natural language processing (NLP) area. Inspired by their effectiveness in facilitating various NLP tasks, more and more researchers attempt to adopt these WE models for their software engineering (SE) tasks, of which semantic representation of software artifacts such as bug reports and code snippets is the basis for further model building. However, existing studies are generally isolated from each other without comprehensive comparison and discussion. This not only makes the best practice of such cross-discipline technique adoption buried in scattered papers, but also makes us kind of blind to current progress in the semantic representation of SE artifacts. To this end, we decided to perform a comprehensive study on the use of WE models in the SE domain. 181 primary studies published in mainstream software engineering venues are collected for analysis. Several research questions related to the SE applications, the training strategy of WE models, the comparison with traditional semantic representation methods, etc., are answered. With the answers, we get a systematical view of the current practice of using WE for the SE domain, and figure out the challenges and actions in adopting or developing practical semantic representation approaches for the SE artifacts used in a series of SE tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17634
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Comprehensive Study on the Use of Word Embedding Models in Software Engineering Domain
Chen, Xiaohan
Zou, Weiqin
Zhi, Lianyi
Meng, Qianshuang
Zhang, Jingxuan
Software Engineering
Word embedding (WE) techniques are advanced textual semantic representation models oriented from the natural language processing (NLP) area. Inspired by their effectiveness in facilitating various NLP tasks, more and more researchers attempt to adopt these WE models for their software engineering (SE) tasks, of which semantic representation of software artifacts such as bug reports and code snippets is the basis for further model building. However, existing studies are generally isolated from each other without comprehensive comparison and discussion. This not only makes the best practice of such cross-discipline technique adoption buried in scattered papers, but also makes us kind of blind to current progress in the semantic representation of SE artifacts. To this end, we decided to perform a comprehensive study on the use of WE models in the SE domain. 181 primary studies published in mainstream software engineering venues are collected for analysis. Several research questions related to the SE applications, the training strategy of WE models, the comparison with traditional semantic representation methods, etc., are answered. With the answers, we get a systematical view of the current practice of using WE for the SE domain, and figure out the challenges and actions in adopting or developing practical semantic representation approaches for the SE artifacts used in a series of SE tasks.
title A Comprehensive Study on the Use of Word Embedding Models in Software Engineering Domain
topic Software Engineering
url https://arxiv.org/abs/2505.17634