CodeT5-RNN: Reinforcing Contextual Embeddings for Enhanced Code Comprehension

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rahman, Md Mostafizer, Shiplu, Ariful Islam, Watanobe, Yutaka, Amin, Md Faizul Ibne, Naqvi, Syed Rameez, Liu, Fang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914416250847232
author Rahman, Md Mostafizer
Shiplu, Ariful Islam
Watanobe, Yutaka
Amin, Md Faizul Ibne
Naqvi, Syed Rameez
Liu, Fang
author_facet Rahman, Md Mostafizer
Shiplu, Ariful Islam
Watanobe, Yutaka
Amin, Md Faizul Ibne
Naqvi, Syed Rameez
Liu, Fang
contents Contextual embeddings generated by LLMs exhibit strong positional inductive biases, which can limit their ability to fully capture long-range, order-sensitive dependencies in highly structured source code. Consequently, how to further refine and enhance LLM embeddings for improved code understanding remains an open research question. To address this gap, we propose a hybrid LLM-RNN framework that reinforces LLM-generated contextual embeddings with a sequential RNN architecture. The embeddings reprocessing step aims to reinforce sequential semantics and strengthen order-aware dependencies inherent in source code. We evaluate the proposed hybrid models on both benchmark and real-world coding datasets. The experimental results show that the RoBERTa-BiGRU and CodeBERT-GRU models achieved accuracies of 66.40% and 66.03%, respectively, on the defect detection benchmark dataset, representing improvements of approximately 5.35% and 3.95% over the standalone RoBERTa and CodeBERT models. Furthermore, the CodeT5-GRU and CodeT5+-BiGRU models achieved accuracies of 67.90% and 67.79%, respectively, surpassing their base models and outperforming RoBERTa-BiGRU and CodeBERT-GRU by a notable margin. In addition, CodeT5-GRU model attains weighted and macro F1-scores of 67.18% and 67.00%, respectively, on the same dataset. Extensive experiments across three real-world datasets further demonstrate consistent and statistically significant improvements over standalone LLMs. Overall, our findings indicate that reprocessing contextual embeddings with RNN architectures enhances code understanding performance in LLM-based models.
format Preprint
id arxiv_https___arxiv_org_abs_2603_17821
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CodeT5-RNN: Reinforcing Contextual Embeddings for Enhanced Code Comprehension
Rahman, Md Mostafizer
Shiplu, Ariful Islam
Watanobe, Yutaka
Amin, Md Faizul Ibne
Naqvi, Syed Rameez
Liu, Fang
Software Engineering
Contextual embeddings generated by LLMs exhibit strong positional inductive biases, which can limit their ability to fully capture long-range, order-sensitive dependencies in highly structured source code. Consequently, how to further refine and enhance LLM embeddings for improved code understanding remains an open research question. To address this gap, we propose a hybrid LLM-RNN framework that reinforces LLM-generated contextual embeddings with a sequential RNN architecture. The embeddings reprocessing step aims to reinforce sequential semantics and strengthen order-aware dependencies inherent in source code. We evaluate the proposed hybrid models on both benchmark and real-world coding datasets. The experimental results show that the RoBERTa-BiGRU and CodeBERT-GRU models achieved accuracies of 66.40% and 66.03%, respectively, on the defect detection benchmark dataset, representing improvements of approximately 5.35% and 3.95% over the standalone RoBERTa and CodeBERT models. Furthermore, the CodeT5-GRU and CodeT5+-BiGRU models achieved accuracies of 67.90% and 67.79%, respectively, surpassing their base models and outperforming RoBERTa-BiGRU and CodeBERT-GRU by a notable margin. In addition, CodeT5-GRU model attains weighted and macro F1-scores of 67.18% and 67.00%, respectively, on the same dataset. Extensive experiments across three real-world datasets further demonstrate consistent and statistically significant improvements over standalone LLMs. Overall, our findings indicate that reprocessing contextual embeddings with RNN architectures enhances code understanding performance in LLM-based models.
title CodeT5-RNN: Reinforcing Contextual Embeddings for Enhanced Code Comprehension
topic Software Engineering
url https://arxiv.org/abs/2603.17821