High-quality data augmentation for code comment classification
Fuente:
arXiv
Saved in:
| Main Authors: | Borsani, Thomas, Rosani, Andrea, Di Fatta, Giuseppe |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimizing Deep Learning Models to Address Class Imbalance in Code Comment Classification
by: Mock, Moritz, et al.
Published: (2025)
by: Mock, Moritz, et al.
Published: (2025)
Gradient Similarity Surgery in Multi-Task Deep Learning
by: Borsani, Thomas, et al.
Published: (2025)
by: Borsani, Thomas, et al.
Published: (2025)
Ensuring Reliability in Programming Knowledge Tracing: A Re-evaluation of Attention-augmented Models and Experimental Protocols
by: Kim, Jaewook, et al.
Published: (2026)
by: Kim, Jaewook, et al.
Published: (2026)
Retrieval-augmented code completion for local projects using large language models
by: Hostnik, Marko, et al.
Published: (2024)
by: Hostnik, Marko, et al.
Published: (2024)
pyAKI -- An Open Source Solution to Automated KDIGO classification
by: Porschen, Christian, et al.
Published: (2024)
by: Porschen, Christian, et al.
Published: (2024)
What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering
by: Errica, Federico, et al.
Published: (2024)
by: Errica, Federico, et al.
Published: (2024)
Practical programming research of Linear DML model based on the simplest Python code: From the standpoint of novice researchers
by: Yao, Shunxin
Published: (2025)
by: Yao, Shunxin
Published: (2025)
David vs. Goliath: A comparative study of different-sized LLMs for code generation in the domain of automotive scenario generation
by: Bauerfeind, Philipp, et al.
Published: (2025)
by: Bauerfeind, Philipp, et al.
Published: (2025)
Are your comments outdated? Towards automatically detecting code-comment consistency
by: Huang, Yuan, et al.
Published: (2024)
by: Huang, Yuan, et al.
Published: (2024)
QualiTagger: Automating software quality detection in issue trackers
by: Shivashankar, Karthik, et al.
Published: (2025)
by: Shivashankar, Karthik, et al.
Published: (2025)
Latent Regularization in Generative Test Input Generation
by: Merabishvili, Giorgi, et al.
Published: (2026)
by: Merabishvili, Giorgi, et al.
Published: (2026)
When simplicity meets effectiveness: Detecting code comments coherence with word embeddings and LSTM
by: Igbomezie, Michael Dubem, et al.
Published: (2024)
by: Igbomezie, Michael Dubem, et al.
Published: (2024)
Engineering Resource-constrained Software Systems with DNN Components: a Concept-based Pruning Approach
by: Formica, Federico, et al.
Published: (2026)
by: Formica, Federico, et al.
Published: (2026)
FRANC: A Lightweight Framework for High-Quality Code Generation
by: Siddiq, Mohammed Latif, et al.
Published: (2023)
by: Siddiq, Mohammed Latif, et al.
Published: (2023)
Data Preparation for Fairness-Performance Trade-Offs: A Practitioner-Friendly Alternative?
by: Voria, Gianmario, et al.
Published: (2024)
by: Voria, Gianmario, et al.
Published: (2024)
HyperNet-Adaptation for Diffusion-Based Test Case Generation
by: Weißl, Oliver, et al.
Published: (2026)
by: Weißl, Oliver, et al.
Published: (2026)
Measuring memorization in RLHF for code completion
by: Pappu, Aneesh, et al.
Published: (2024)
by: Pappu, Aneesh, et al.
Published: (2024)
CodeIF: Benchmarking the Instruction-Following Capabilities of Large Language Models for Code Generation
by: Yan, Kaiwen, et al.
Published: (2025)
by: Yan, Kaiwen, et al.
Published: (2025)
Reinforcement Learning from Automatic Feedback for High-Quality Unit Test Generation
by: Steenhoek, Benjamin, et al.
Published: (2023)
by: Steenhoek, Benjamin, et al.
Published: (2023)
Reinforcement Learning from Automatic Feedback for High-Quality Unit Test Generation
by: Steenhoek, Benjamin, et al.
Published: (2024)
by: Steenhoek, Benjamin, et al.
Published: (2024)
NeSy is alive and well: A LLM-driven symbolic approach for better code comment data generation and classification
by: Akl, Hanna Abi
Published: (2024)
by: Akl, Hanna Abi
Published: (2024)
Breaking the Silence: the Threats of Using LLMs in Software Engineering
by: Sallou, June, et al.
Published: (2023)
by: Sallou, June, et al.
Published: (2023)
Aggregating empirical evidence from data strategy studies: a case on model quantization
by: del Rey, Santiago, et al.
Published: (2025)
by: del Rey, Santiago, et al.
Published: (2025)
Overcoming linguistic barriers in code assistants: creating a QLoRA adapter to improve support for Russian-language code writing instructions
by: Pronin, C. B., et al.
Published: (2024)
by: Pronin, C. B., et al.
Published: (2024)
An Empirical Analysis of Machine Learning Model and Dataset Documentation, Supply Chain, and Licensing Challenges on Hugging Face
by: Stalnaker, Trevor, et al.
Published: (2025)
by: Stalnaker, Trevor, et al.
Published: (2025)
Targeted Deep Learning System Boundary Testing
by: Weißl, Oliver, et al.
Published: (2024)
by: Weißl, Oliver, et al.
Published: (2024)
Fault Localization via Fine-tuning Large Language Models with Mutation Generated Stack Traces
by: Jambigi, Neetha, et al.
Published: (2025)
by: Jambigi, Neetha, et al.
Published: (2025)
"You still have to study" -- On the Security of LLM generated code
by: Goetz, Stefan, et al.
Published: (2024)
by: Goetz, Stefan, et al.
Published: (2024)
A Catalog of Fairness-Aware Practices in Machine Learning Engineering
by: Voria, Gianmario, et al.
Published: (2024)
by: Voria, Gianmario, et al.
Published: (2024)
Beyond the Comfort Zone: Emerging Solutions to Overcome Challenges in Integrating LLMs into Software Products
by: Nahar, Nadia, et al.
Published: (2024)
by: Nahar, Nadia, et al.
Published: (2024)
SpIDER: Spatially Informed Dense Embedding Retrieval for Software Issue Localization
by: Chaudhari, Shravan, et al.
Published: (2025)
by: Chaudhari, Shravan, et al.
Published: (2025)
QiMeng-PRepair: Precise Code Repair via Edit-Aware Reward Optimization
by: Ke, Changxin, et al.
Published: (2026)
by: Ke, Changxin, et al.
Published: (2026)
OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents
by: Kuntz, Thomas, et al.
Published: (2025)
by: Kuntz, Thomas, et al.
Published: (2025)
Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks
by: Tao, Hongyuan, et al.
Published: (2025)
by: Tao, Hongyuan, et al.
Published: (2025)
Automating Benchmark Design
by: Dsouza, Amanda, et al.
Published: (2025)
by: Dsouza, Amanda, et al.
Published: (2025)
Ontology- and LLM-based Data Harmonization for Federated Learning in Healthcare
by: Kokash, Natallia, et al.
Published: (2025)
by: Kokash, Natallia, et al.
Published: (2025)
Mining patterns in syntax trees to automate code reviews of student solutions for programming exercises
by: Van Petegem, Charlotte, et al.
Published: (2024)
by: Van Petegem, Charlotte, et al.
Published: (2024)
Predicting Safety Misbehaviours in Autonomous Driving Systems using Uncertainty Quantification
by: Grewal, Ruben, et al.
Published: (2024)
by: Grewal, Ruben, et al.
Published: (2024)
Insights into resource utilization of code small language models serving with runtime engines and execution providers
by: Durán, Francisco, et al.
Published: (2024)
by: Durán, Francisco, et al.
Published: (2024)
evomap: A Toolbox for Dynamic Mapping in Python
by: Matthe, Maximilian
Published: (2025)
by: Matthe, Maximilian
Published: (2025)
Similar Items
-
Optimizing Deep Learning Models to Address Class Imbalance in Code Comment Classification
by: Mock, Moritz, et al.
Published: (2025) -
Gradient Similarity Surgery in Multi-Task Deep Learning
by: Borsani, Thomas, et al.
Published: (2025) -
Ensuring Reliability in Programming Knowledge Tracing: A Re-evaluation of Attention-augmented Models and Experimental Protocols
by: Kim, Jaewook, et al.
Published: (2026) -
Retrieval-augmented code completion for local projects using large language models
by: Hostnik, Marko, et al.
Published: (2024) -
pyAKI -- An Open Source Solution to Automated KDIGO classification
by: Porschen, Christian, et al.
Published: (2024)