Toward a Theory of Causation for Interpreting Neural Code Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Palacio, David N., Velasco, Alejandro, Cooper, Nathan, Rodriguez, Alvaro, Moran, Kevin, Poshyvanyk, Denys
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916180733722624
author Palacio, David N.
Velasco, Alejandro
Cooper, Nathan
Rodriguez, Alvaro
Moran, Kevin
Poshyvanyk, Denys
author_facet Palacio, David N.
Velasco, Alejandro
Cooper, Nathan
Rodriguez, Alvaro
Moran, Kevin
Poshyvanyk, Denys
contents Neural Language Models of Code, or Neural Code Models (NCMs), are rapidly progressing from research prototypes to commercial developer tools. As such, understanding the capabilities and limitations of such models is becoming critical. However, the abilities of these models are typically measured using automated metrics that often only reveal a portion of their real-world performance. While, in general, the performance of NCMs appears promising, currently much is unknown about how such models arrive at decisions. To this end, this paper introduces $do_{code}$, a post hoc interpretability method specific to NCMs that is capable of explaining model predictions. $do_{code}$ is based upon causal inference to enable programming language-oriented explanations. While the theoretical underpinnings of $do_{code}$ are extensible to exploring different model properties, we provide a concrete instantiation that aims to mitigate the impact of spurious correlations by grounding explanations of model behavior in properties of programming languages. To demonstrate the practical benefit of $do_{code}$, we illustrate the insights that our framework can provide by performing a case study on two popular deep learning architectures and ten NCMs. The results of this case study illustrate that our studied NCMs are sensitive to changes in code syntax. All our NCMs, except for the BERT-like model, statistically learn to predict tokens related to blocks of code (\eg brackets, parenthesis, semicolon) with less confounding bias as compared to other programming language constructs. These insights demonstrate the potential of $do_{code}$ as a useful method to detect and facilitate the elimination of confounding bias in NCMs.
format Preprint
id arxiv_https___arxiv_org_abs_2302_03788
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Toward a Theory of Causation for Interpreting Neural Code Models
Palacio, David N.
Velasco, Alejandro
Cooper, Nathan
Rodriguez, Alvaro
Moran, Kevin
Poshyvanyk, Denys
Software Engineering
Artificial Intelligence
Machine Learning
Methodology
Neural Language Models of Code, or Neural Code Models (NCMs), are rapidly progressing from research prototypes to commercial developer tools. As such, understanding the capabilities and limitations of such models is becoming critical. However, the abilities of these models are typically measured using automated metrics that often only reveal a portion of their real-world performance. While, in general, the performance of NCMs appears promising, currently much is unknown about how such models arrive at decisions. To this end, this paper introduces $do_{code}$, a post hoc interpretability method specific to NCMs that is capable of explaining model predictions. $do_{code}$ is based upon causal inference to enable programming language-oriented explanations. While the theoretical underpinnings of $do_{code}$ are extensible to exploring different model properties, we provide a concrete instantiation that aims to mitigate the impact of spurious correlations by grounding explanations of model behavior in properties of programming languages. To demonstrate the practical benefit of $do_{code}$, we illustrate the insights that our framework can provide by performing a case study on two popular deep learning architectures and ten NCMs. The results of this case study illustrate that our studied NCMs are sensitive to changes in code syntax. All our NCMs, except for the BERT-like model, statistically learn to predict tokens related to blocks of code (\eg brackets, parenthesis, semicolon) with less confounding bias as compared to other programming language constructs. These insights demonstrate the potential of $do_{code}$ as a useful method to detect and facilitate the elimination of confounding bias in NCMs.
title Toward a Theory of Causation for Interpreting Neural Code Models
topic Software Engineering
Artificial Intelligence
Machine Learning
Methodology
url https://arxiv.org/abs/2302.03788