StRuCom: A Novel Dataset of Structured Code Comments in Russian

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dziuba, Maria, Malykh, Valentin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912379569176576
author Dziuba, Maria
Malykh, Valentin
author_facet Dziuba, Maria
Malykh, Valentin
contents Structured code comments in docstring format are essential for code comprehension and maintenance, but existing machine learning models for their generation perform poorly for Russian compared to English. To bridge this gap, we present StRuCom - the first large-scale dataset (153K examples) specifically designed for Russian code documentation. Unlike machine-translated English datasets that distort terminology (e.g., technical loanwords vs. literal translations) and docstring structures, StRuCom combines human-written comments from Russian GitHub repositories with synthetically generated ones, ensuring compliance with Python, Java, JavaScript, C#, and Go standards through automated validation. Fine-tuning Qwen2.5-Coder models (0.5B-7B) on StRuCom shows statistically significant improvements of chrf++ and BERTScore over baseline models.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11026
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle StRuCom: A Novel Dataset of Structured Code Comments in Russian
Dziuba, Maria
Malykh, Valentin
Computation and Language
Artificial Intelligence
Machine Learning
Software Engineering
Structured code comments in docstring format are essential for code comprehension and maintenance, but existing machine learning models for their generation perform poorly for Russian compared to English. To bridge this gap, we present StRuCom - the first large-scale dataset (153K examples) specifically designed for Russian code documentation. Unlike machine-translated English datasets that distort terminology (e.g., technical loanwords vs. literal translations) and docstring structures, StRuCom combines human-written comments from Russian GitHub repositories with synthetically generated ones, ensuring compliance with Python, Java, JavaScript, C#, and Go standards through automated validation. Fine-tuning Qwen2.5-Coder models (0.5B-7B) on StRuCom shows statistically significant improvements of chrf++ and BERTScore over baseline models.
title StRuCom: A Novel Dataset of Structured Code Comments in Russian
topic Computation and Language
Artificial Intelligence
Machine Learning
Software Engineering
url https://arxiv.org/abs/2505.11026