Advancing Language Models for Code-related Tasks

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autor principal: Tian, Zhao
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912809724411904
author Tian, Zhao
author_facet Tian, Zhao
contents Recent advances in language models (LMs) have driven significant progress in various software engineering tasks. However, existing LMs still struggle with complex programming scenarios due to limitations in data quality, model architecture, and reasoning capability. This research systematically addresses these challenges through three complementary directions: (1) improving code data quality with a code difference-guided adversarial augmentation technique (CODA) and a code denoising technique (CodeDenoise); (2) enhancing model architecture via syntax-guided code LMs (LEAM and LEAM++); and (3) advancing model reasoning with a prompting technique (muFiX) and an agent-based technique (Specine). These techniques aim to promote the practical adoption of LMs in software development and further advance intelligent software engineering.
format Preprint
id arxiv_https___arxiv_org_abs_2601_04526
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Advancing Language Models for Code-related Tasks
Tian, Zhao
Software Engineering
Artificial Intelligence
Computation and Language
Recent advances in language models (LMs) have driven significant progress in various software engineering tasks. However, existing LMs still struggle with complex programming scenarios due to limitations in data quality, model architecture, and reasoning capability. This research systematically addresses these challenges through three complementary directions: (1) improving code data quality with a code difference-guided adversarial augmentation technique (CODA) and a code denoising technique (CodeDenoise); (2) enhancing model architecture via syntax-guided code LMs (LEAM and LEAM++); and (3) advancing model reasoning with a prompting technique (muFiX) and an agent-based technique (Specine). These techniques aim to promote the practical adoption of LMs in software development and further advance intelligent software engineering.
title Advancing Language Models for Code-related Tasks
topic Software Engineering
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2601.04526