Do Large Language Models Pay Similar Attention Like Human Programmers When Generating Code?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kou, Bonan, Chen, Shengmai, Wang, Zhijie, Ma, Lei, Zhang, Tianyi
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929352948580352
author Kou, Bonan
Chen, Shengmai
Wang, Zhijie
Ma, Lei
Zhang, Tianyi
author_facet Kou, Bonan
Chen, Shengmai
Wang, Zhijie
Ma, Lei
Zhang, Tianyi
contents Large Language Models (LLMs) have recently been widely used for code generation. Due to the complexity and opacity of LLMs, little is known about how these models generate code. We made the first attempt to bridge this knowledge gap by investigating whether LLMs attend to the same parts of a task description as human programmers during code generation. An analysis of six LLMs, including GPT-4, on two popular code generation benchmarks revealed a consistent misalignment between LLMs' and programmers' attention. We manually analyzed 211 incorrect code snippets and found five attention patterns that can be used to explain many code generation errors. Finally, a user study showed that model attention computed by a perturbation-based method is often favored by human programmers. Our findings highlight the need for human-aligned LLMs for better interpretability and programmer trust.
format Preprint
id arxiv_https___arxiv_org_abs_2306_01220
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Do Large Language Models Pay Similar Attention Like Human Programmers When Generating Code?
Kou, Bonan
Chen, Shengmai
Wang, Zhijie
Ma, Lei
Zhang, Tianyi
Software Engineering
Human-Computer Interaction
Machine Learning
Large Language Models (LLMs) have recently been widely used for code generation. Due to the complexity and opacity of LLMs, little is known about how these models generate code. We made the first attempt to bridge this knowledge gap by investigating whether LLMs attend to the same parts of a task description as human programmers during code generation. An analysis of six LLMs, including GPT-4, on two popular code generation benchmarks revealed a consistent misalignment between LLMs' and programmers' attention. We manually analyzed 211 incorrect code snippets and found five attention patterns that can be used to explain many code generation errors. Finally, a user study showed that model attention computed by a perturbation-based method is often favored by human programmers. Our findings highlight the need for human-aligned LLMs for better interpretability and programmer trust.
title Do Large Language Models Pay Similar Attention Like Human Programmers When Generating Code?
topic Software Engineering
Human-Computer Interaction
Machine Learning
url https://arxiv.org/abs/2306.01220