How do Humans and LLMs Process Confusing Code?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Abdelsalam, Youssef, Peitek, Norman, Maurer, Anna-Maria, Toneva, Mariya, Apel, Sven
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908503723999232
author Abdelsalam, Youssef
Peitek, Norman
Maurer, Anna-Maria
Toneva, Mariya
Apel, Sven
author_facet Abdelsalam, Youssef
Peitek, Norman
Maurer, Anna-Maria
Toneva, Mariya
Apel, Sven
contents Already today, humans and programming assistants based on large language models (LLMs) collaborate in everyday programming tasks. Clearly, a misalignment between how LLMs and programmers comprehend code can lead to misunderstandings, inefficiencies, low code quality, and bugs. A key question in this space is whether humans and LLMs are confused by the same kind of code. This would not only guide our choices of integrating LLMs in software engineering workflows, but also inform about possible improvements of LLMs. To this end, we conducted an empirical study comparing an LLM to human programmers comprehending clean and confusing code. We operationalized comprehension for the LLM by using LLM perplexity, and for human programmers using neurophysiological responses (in particular, EEG-based fixation-related potentials). We found that LLM perplexity spikes correlate both in terms of location and amplitude with human neurophysiological responses that indicate confusion. This result suggests that LLMs and humans are similarly confused about the code. Based on these findings, we devised a data-driven, LLM-based approach to identify regions of confusion in code that elicit confusion in human programmers.
format Preprint
id arxiv_https___arxiv_org_abs_2508_18547
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle How do Humans and LLMs Process Confusing Code?
Abdelsalam, Youssef
Peitek, Norman
Maurer, Anna-Maria
Toneva, Mariya
Apel, Sven
Software Engineering
Already today, humans and programming assistants based on large language models (LLMs) collaborate in everyday programming tasks. Clearly, a misalignment between how LLMs and programmers comprehend code can lead to misunderstandings, inefficiencies, low code quality, and bugs. A key question in this space is whether humans and LLMs are confused by the same kind of code. This would not only guide our choices of integrating LLMs in software engineering workflows, but also inform about possible improvements of LLMs. To this end, we conducted an empirical study comparing an LLM to human programmers comprehending clean and confusing code. We operationalized comprehension for the LLM by using LLM perplexity, and for human programmers using neurophysiological responses (in particular, EEG-based fixation-related potentials). We found that LLM perplexity spikes correlate both in terms of location and amplitude with human neurophysiological responses that indicate confusion. This result suggests that LLMs and humans are similarly confused about the code. Based on these findings, we devised a data-driven, LLM-based approach to identify regions of confusion in code that elicit confusion in human programmers.
title How do Humans and LLMs Process Confusing Code?
topic Software Engineering
url https://arxiv.org/abs/2508.18547