Design, Implementation and Evaluation of a Novel Programming Language Topic Classification Workflow

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Michael, Tian, Yuan, Guizani, Mariam
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914055761952768
author Zhang, Michael
Tian, Yuan
Guizani, Mariam
author_facet Zhang, Michael
Tian, Yuan
Guizani, Mariam
contents As software systems grow in scale and complexity, understanding the distribution of programming language topics within source code becomes increasingly important for guiding technical decisions, improving onboarding, and informing tooling and education. This paper presents the design, implementation, and evaluation of a novel programming language topic classification workflow. Our approach combines a multi-label Support Vector Machine (SVM) with a sliding window and voting strategy to enable fine-grained localization of core language concepts such as operator overloading, virtual functions, inheritance, and templates. Trained on the IBM Project CodeNet dataset, our model achieves an average F1 score of 0.90 across topics and 0.75 in code-topic highlight. Our findings contribute empirical insights and a reusable pipeline for researchers and practitioners interested in code analysis and data-driven software engineering.
format Preprint
id arxiv_https___arxiv_org_abs_2509_20631
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Design, Implementation and Evaluation of a Novel Programming Language Topic Classification Workflow
Zhang, Michael
Tian, Yuan
Guizani, Mariam
Software Engineering
Machine Learning
As software systems grow in scale and complexity, understanding the distribution of programming language topics within source code becomes increasingly important for guiding technical decisions, improving onboarding, and informing tooling and education. This paper presents the design, implementation, and evaluation of a novel programming language topic classification workflow. Our approach combines a multi-label Support Vector Machine (SVM) with a sliding window and voting strategy to enable fine-grained localization of core language concepts such as operator overloading, virtual functions, inheritance, and templates. Trained on the IBM Project CodeNet dataset, our model achieves an average F1 score of 0.90 across topics and 0.75 in code-topic highlight. Our findings contribute empirical insights and a reusable pipeline for researchers and practitioners interested in code analysis and data-driven software engineering.
title Design, Implementation and Evaluation of a Novel Programming Language Topic Classification Workflow
topic Software Engineering
Machine Learning
url https://arxiv.org/abs/2509.20631