CodeCrash: Exposing LLM Fragility to Misleading Natural Language in Code Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Lam, Man Ho, Wang, Chaozheng, Huang, Jen-tse, Lyu, Michael R. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades
by: Lam, Man Ho, et al.
Published: (2026)
by: Lam, Man Ho, et al.
Published: (2026)
TREAT: A Code LLMs Trustworthiness / Reliability Evaluation and Testing Framework
by: Gao, Shuzheng, et al.
Published: (2025)
by: Gao, Shuzheng, et al.
Published: (2025)
DesignBench: A Comprehensive Benchmark for MLLM-based Front-end Code Generation
by: Xiao, Jingyu, et al.
Published: (2025)
by: Xiao, Jingyu, et al.
Published: (2025)
Learning to Ask: When LLM Agents Meet Unclear Instruction
by: Wang, Wenxuan, et al.
Published: (2024)
by: Wang, Wenxuan, et al.
Published: (2024)
Automatically Generating UI Code from Screenshot: A Divide-and-Conquer-Based Approach
by: Wan, Yuxuan, et al.
Published: (2024)
by: Wan, Yuxuan, et al.
Published: (2024)
Constraint Decay: The Fragility of LLM Agents in Backend Code Generation
by: Dente, Francesco, et al.
Published: (2026)
by: Dente, Francesco, et al.
Published: (2026)
ComUICoder: Component-based Reusable UI Code Generation for Complex Websites via Semantic Segmentation and Element-wise Feedback
by: Xiao, Jingyu, et al.
Published: (2026)
by: Xiao, Jingyu, et al.
Published: (2026)
From Laboratory to Real-World Applications: Benchmarking Agentic Code Reasoning at the Repository Level
by: Li, Jia, et al.
Published: (2026)
by: Li, Jia, et al.
Published: (2026)
Revisiting the Role of Natural Language Code Comments in Code Translation
by: Gupta, Monika, et al.
Published: (2026)
by: Gupta, Monika, et al.
Published: (2026)
Enhancing LLM-Based Coding Tools through Native Integration of IDE-Derived Static Context
by: Li, Yichen, et al.
Published: (2024)
by: Li, Yichen, et al.
Published: (2024)
CodeCoT: Tackling Code Syntax Errors in CoT Reasoning for Code Generation
by: Huang, Dong, et al.
Published: (2023)
by: Huang, Dong, et al.
Published: (2023)
CodeMMR: Bridging Natural Language, Code, and Image for Unified Retrieval
by: Geng, Jiahui, et al.
Published: (2026)
by: Geng, Jiahui, et al.
Published: (2026)
LogicAsker: Evaluating and Improving the Logical Reasoning Ability of Large Language Models
by: Wan, Yuxuan, et al.
Published: (2024)
by: Wan, Yuxuan, et al.
Published: (2024)
90% Faster, 100% Code-Free: MLLM-Driven Zero-Code 3D Game Development
by: Yang, Runxin, et al.
Published: (2025)
by: Yang, Runxin, et al.
Published: (2025)
SPENCER: Self-Adaptive Model Distillation for Efficient Code Retrieval
by: Gu, Wenchao, et al.
Published: (2025)
by: Gu, Wenchao, et al.
Published: (2025)
SEER: Enhancing Chain-of-Thought Code Generation through Self-Exploring Deep Reasoning
by: Gao, Shuzheng, et al.
Published: (2025)
by: Gao, Shuzheng, et al.
Published: (2025)
FastCode: Fast and Cost-Efficient Code Understanding and Reasoning
by: Li, Zhonghang, et al.
Published: (2026)
by: Li, Zhonghang, et al.
Published: (2026)
Where Code Meets Natural Language: Taxonomy-Driven Information Flow Analysis for LLM-Integrated Applications
by: Xu, Zihao, et al.
Published: (2026)
by: Xu, Zihao, et al.
Published: (2026)
Class-Level Code Generation from Natural Language Using Iterative, Tool-Enhanced Reasoning over Repository
by: Deshpande, Ajinkya, et al.
Published: (2024)
by: Deshpande, Ajinkya, et al.
Published: (2024)
CLAP: Learning Transferable Binary Code Representations with Natural Language Supervision
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
LLM4EFFI: Leveraging Large Language Models to Enhance Code Efficiency and Correctness
by: Ye, Tong, et al.
Published: (2025)
by: Ye, Tong, et al.
Published: (2025)
CodeS: Natural Language to Code Repository via Multi-Layer Sketch
by: Zan, Daoguang, et al.
Published: (2024)
by: Zan, Daoguang, et al.
Published: (2024)
Uncovering LLM-Generated Code: A Zero-Shot Synthetic Code Detector via Code Rewriting
by: Ye, Tong, et al.
Published: (2024)
by: Ye, Tong, et al.
Published: (2024)
Learn to Code Sustainably: An Empirical Study on LLM-based Green Code Generation
by: Vartziotis, Tina, et al.
Published: (2024)
by: Vartziotis, Tina, et al.
Published: (2024)
Code Copycat Conundrum: Demystifying Repetition in LLM-based Code Generation
by: Liu, Mingwei, et al.
Published: (2025)
by: Liu, Mingwei, et al.
Published: (2025)
An Execution-Verified Multi-Language Benchmark for Code Semantic Reasoning
by: Li, Yikun, et al.
Published: (2026)
by: Li, Yikun, et al.
Published: (2026)
Identifying the Achilles' Heel: An Iterative Method for Dynamically Uncovering Factual Errors in Large Language Models
by: Wang, Wenxuan, et al.
Published: (2024)
by: Wang, Wenxuan, et al.
Published: (2024)
From Natural Language to Verified Code: Toward AI Assisted Problem-to-Code Generation with Dafny-Based Formal Verification
by: Erfan, Md, et al.
Published: (2026)
by: Erfan, Md, et al.
Published: (2026)
Top Pass: Improve Code Generation by Pass@k-Maximized Code Ranking
by: Lyu, Zhi-Cun, et al.
Published: (2024)
by: Lyu, Zhi-Cun, et al.
Published: (2024)
Automated Code Review Using Large Language Models with Symbolic Reasoning
by: Icoz, Busra, et al.
Published: (2025)
by: Icoz, Busra, et al.
Published: (2025)
CRPE: Expanding The Reasoning Capability of Large Language Model for Code Generation
by: Gui, Ningxin, et al.
Published: (2025)
by: Gui, Ningxin, et al.
Published: (2025)
Pragmatic Reasoning improves LLM Code Generation
by: Cao, Zhuchen, et al.
Published: (2025)
by: Cao, Zhuchen, et al.
Published: (2025)
WebCompass: Towards Multimodal Web Coding Evaluation for Code Language Models
by: Lei, Xinping, et al.
Published: (2026)
by: Lei, Xinping, et al.
Published: (2026)
Impact of Comments on LLM Comprehension of Legacy Code
by: Sabetto, Rock, et al.
Published: (2025)
by: Sabetto, Rock, et al.
Published: (2025)
CodeSense: a Real-World Benchmark and Dataset for Code Semantic Reasoning
by: Roy, Monoshi Kumar, et al.
Published: (2025)
by: Roy, Monoshi Kumar, et al.
Published: (2025)
Hallucinations in Code Change to Natural Language Generation: Prevalence and Evaluation of Detection Metrics
by: Liu, Chunhua, et al.
Published: (2025)
by: Liu, Chunhua, et al.
Published: (2025)
Uncovering Systematic Failures of LLMs in Verifying Code Against Natural Language Specifications
by: Jin, Haolin, et al.
Published: (2025)
by: Jin, Haolin, et al.
Published: (2025)
Bias Testing and Mitigation in LLM-based Code Generation
by: Huang, Dong, et al.
Published: (2023)
by: Huang, Dong, et al.
Published: (2023)
Position Paper: Programming Language Techniques for Bridging LLM Code Generation Semantic Gaps
by: Du, Yalong, et al.
Published: (2025)
by: Du, Yalong, et al.
Published: (2025)
CodeSift: An LLM-Based Reference-Less Framework for Automatic Code Validation
by: Aggarwal, Pooja, et al.
Published: (2024)
by: Aggarwal, Pooja, et al.
Published: (2024)
Similar Items
-
SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades
by: Lam, Man Ho, et al.
Published: (2026) -
TREAT: A Code LLMs Trustworthiness / Reliability Evaluation and Testing Framework
by: Gao, Shuzheng, et al.
Published: (2025) -
DesignBench: A Comprehensive Benchmark for MLLM-based Front-end Code Generation
by: Xiao, Jingyu, et al.
Published: (2025) -
Learning to Ask: When LLM Agents Meet Unclear Instruction
by: Wang, Wenxuan, et al.
Published: (2024) -
Automatically Generating UI Code from Screenshot: A Divide-and-Conquer-Based Approach
by: Wan, Yuxuan, et al.
Published: (2024)