CodeCrash: Exposing LLM Fragility to Misleading Natural Language in Code Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lam, Man Ho, Wang, Chaozheng, Huang, Jen-tse, Lyu, Michael R.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912641715273728
author Lam, Man Ho
Wang, Chaozheng
Huang, Jen-tse
Lyu, Michael R.
author_facet Lam, Man Ho
Wang, Chaozheng
Huang, Jen-tse
Lyu, Michael R.
contents Large Language Models (LLMs) have recently demonstrated strong capabilities in code-related tasks, but their robustness in code reasoning under perturbations remains underexplored. We introduce CodeCrash, a stress-testing framework with 1,279 questions from CruxEval and LiveCodeBench, designed to evaluate reasoning reliability under structural perturbations and misleading natural language (NL) contexts. Through a systematic evaluation of 17 LLMs, we find that models often shortcut reasoning by over-relying on NL cues, leading to an average performance degradation of 23.2% in output prediction tasks. Even with Chain-of-Thought reasoning, models on average still have a 13.8% drop due to distractibility and rationalization, revealing a lack of critical reasoning capability to distinguish the actual code behaviors. While Large Reasoning Models with internal reasoning mechanisms improve robustness by fostering critical thinking, plausible yet incorrect hints can trigger pathological self-reflection, causing 2-3 times token consumption and even catastrophic cognitive dissonance in extreme cases for QwQ-32B. We refer to this phenomenon as Reasoning Collapse. CodeCrash provides a rigorous benchmark for evaluating robustness in code reasoning, guiding future research and development toward more reliable and resilient models.
format Preprint
id arxiv_https___arxiv_org_abs_2504_14119
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CodeCrash: Exposing LLM Fragility to Misleading Natural Language in Code Reasoning
Lam, Man Ho
Wang, Chaozheng
Huang, Jen-tse
Lyu, Michael R.
Artificial Intelligence
Software Engineering
Large Language Models (LLMs) have recently demonstrated strong capabilities in code-related tasks, but their robustness in code reasoning under perturbations remains underexplored. We introduce CodeCrash, a stress-testing framework with 1,279 questions from CruxEval and LiveCodeBench, designed to evaluate reasoning reliability under structural perturbations and misleading natural language (NL) contexts. Through a systematic evaluation of 17 LLMs, we find that models often shortcut reasoning by over-relying on NL cues, leading to an average performance degradation of 23.2% in output prediction tasks. Even with Chain-of-Thought reasoning, models on average still have a 13.8% drop due to distractibility and rationalization, revealing a lack of critical reasoning capability to distinguish the actual code behaviors. While Large Reasoning Models with internal reasoning mechanisms improve robustness by fostering critical thinking, plausible yet incorrect hints can trigger pathological self-reflection, causing 2-3 times token consumption and even catastrophic cognitive dissonance in extreme cases for QwQ-32B. We refer to this phenomenon as Reasoning Collapse. CodeCrash provides a rigorous benchmark for evaluating robustness in code reasoning, guiding future research and development toward more reliable and resilient models.
title CodeCrash: Exposing LLM Fragility to Misleading Natural Language in Code Reasoning
topic Artificial Intelligence
Software Engineering
url https://arxiv.org/abs/2504.14119