A Critical Study of What Code-LLMs (Do Not) Learn

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Anand, Abhinav, Verma, Shweta, Narasimhan, Krishna, Mezini, Mira
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910491711897600
author Anand, Abhinav
Verma, Shweta
Narasimhan, Krishna
Mezini, Mira
author_facet Anand, Abhinav
Verma, Shweta
Narasimhan, Krishna
Mezini, Mira
contents Large Language Models trained on code corpora (code-LLMs) have demonstrated impressive performance in various coding assistance tasks. However, despite their increased size and training dataset, code-LLMs still have limitations such as suggesting codes with syntactic errors, variable misuse etc. Some studies argue that code-LLMs perform well on coding tasks because they use self-attention and hidden representations to encode relations among input tokens. However, previous works have not studied what code properties are not encoded by code-LLMs. In this paper, we conduct a fine-grained analysis of attention maps and hidden representations of code-LLMs. Our study indicates that code-LLMs only encode relations among specific subsets of input tokens. Specifically, by categorizing input tokens into syntactic tokens and identifiers, we found that models encode relations among syntactic tokens and among identifiers, but they fail to encode relations between syntactic tokens and identifiers. We also found that fine-tuned models encode these relations poorly compared to their pre-trained counterparts. Additionally, larger models with billions of parameters encode significantly less information about code than models with only a few hundred million parameters.
format Preprint
id arxiv_https___arxiv_org_abs_2406_11930
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Critical Study of What Code-LLMs (Do Not) Learn
Anand, Abhinav
Verma, Shweta
Narasimhan, Krishna
Mezini, Mira
Software Engineering
Artificial Intelligence
Computation and Language
Large Language Models trained on code corpora (code-LLMs) have demonstrated impressive performance in various coding assistance tasks. However, despite their increased size and training dataset, code-LLMs still have limitations such as suggesting codes with syntactic errors, variable misuse etc. Some studies argue that code-LLMs perform well on coding tasks because they use self-attention and hidden representations to encode relations among input tokens. However, previous works have not studied what code properties are not encoded by code-LLMs. In this paper, we conduct a fine-grained analysis of attention maps and hidden representations of code-LLMs. Our study indicates that code-LLMs only encode relations among specific subsets of input tokens. Specifically, by categorizing input tokens into syntactic tokens and identifiers, we found that models encode relations among syntactic tokens and among identifiers, but they fail to encode relations between syntactic tokens and identifiers. We also found that fine-tuned models encode these relations poorly compared to their pre-trained counterparts. Additionally, larger models with billions of parameters encode significantly less information about code than models with only a few hundred million parameters.
title A Critical Study of What Code-LLMs (Do Not) Learn
topic Software Engineering
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2406.11930