Design, Implementation and Evaluation of a Novel Programming Language Topic Classification Workflow
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914055761952768 |
|---|---|
| author | Zhang, Michael Tian, Yuan Guizani, Mariam |
| author_facet | Zhang, Michael Tian, Yuan Guizani, Mariam |
| contents | As software systems grow in scale and complexity, understanding the distribution of programming language topics within source code becomes increasingly important for guiding technical decisions, improving onboarding, and informing tooling and education. This paper presents the design, implementation, and evaluation of a novel programming language topic classification workflow. Our approach combines a multi-label Support Vector Machine (SVM) with a sliding window and voting strategy to enable fine-grained localization of core language concepts such as operator overloading, virtual functions, inheritance, and templates. Trained on the IBM Project CodeNet dataset, our model achieves an average F1 score of 0.90 across topics and 0.75 in code-topic highlight. Our findings contribute empirical insights and a reusable pipeline for researchers and practitioners interested in code analysis and data-driven software engineering. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_20631 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Design, Implementation and Evaluation of a Novel Programming Language Topic Classification Workflow Zhang, Michael Tian, Yuan Guizani, Mariam Software Engineering Machine Learning As software systems grow in scale and complexity, understanding the distribution of programming language topics within source code becomes increasingly important for guiding technical decisions, improving onboarding, and informing tooling and education. This paper presents the design, implementation, and evaluation of a novel programming language topic classification workflow. Our approach combines a multi-label Support Vector Machine (SVM) with a sliding window and voting strategy to enable fine-grained localization of core language concepts such as operator overloading, virtual functions, inheritance, and templates. Trained on the IBM Project CodeNet dataset, our model achieves an average F1 score of 0.90 across topics and 0.75 in code-topic highlight. Our findings contribute empirical insights and a reusable pipeline for researchers and practitioners interested in code analysis and data-driven software engineering. |
| title | Design, Implementation and Evaluation of a Novel Programming Language Topic Classification Workflow |
| topic | Software Engineering Machine Learning |
| url | https://arxiv.org/abs/2509.20631 |