Trust Me, I Know This Function: Hijacking LLM Static Analysis using Bias

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bernstein, Shir, Beste, David, Ayzenshteyn, Daniel, Schonherr, Lea, Mirsky, Yisroel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918253438173184
author Bernstein, Shir
Beste, David
Ayzenshteyn, Daniel
Schonherr, Lea
Mirsky, Yisroel
author_facet Bernstein, Shir
Beste, David
Ayzenshteyn, Daniel
Schonherr, Lea
Mirsky, Yisroel
contents Large Language Models (LLMs) are increasingly trusted to perform automated code review and static analysis at scale, supporting tasks such as vulnerability detection, summarization, and refactoring. In this paper, we identify and exploit a critical vulnerability in LLM-based code analysis: an abstraction bias that causes models to overgeneralize familiar programming patterns and overlook small, meaningful bugs. Adversaries can exploit this blind spot to hijack the control flow of the LLM's interpretation with minimal edits and without affecting actual runtime behavior. We refer to this attack as a Familiar Pattern Attack (FPA). We develop a fully automated, black-box algorithm that discovers and injects FPAs into target code. Our evaluation shows that FPAs are not only effective against basic and reasoning models, but are also transferable across model families (OpenAI, Anthropic, Google), and universal across programming languages (Python, C, Rust, Go). Moreover, FPAs remain effective even when models are explicitly warned about the attack via robust system prompts. Finally, we explore positive, defensive uses of FPAs and discuss their broader implications for the reliability and safety of code-oriented LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2508_17361
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Trust Me, I Know This Function: Hijacking LLM Static Analysis using Bias
Bernstein, Shir
Beste, David
Ayzenshteyn, Daniel
Schonherr, Lea
Mirsky, Yisroel
Machine Learning
Cryptography and Security
Large Language Models (LLMs) are increasingly trusted to perform automated code review and static analysis at scale, supporting tasks such as vulnerability detection, summarization, and refactoring. In this paper, we identify and exploit a critical vulnerability in LLM-based code analysis: an abstraction bias that causes models to overgeneralize familiar programming patterns and overlook small, meaningful bugs. Adversaries can exploit this blind spot to hijack the control flow of the LLM's interpretation with minimal edits and without affecting actual runtime behavior. We refer to this attack as a Familiar Pattern Attack (FPA). We develop a fully automated, black-box algorithm that discovers and injects FPAs into target code. Our evaluation shows that FPAs are not only effective against basic and reasoning models, but are also transferable across model families (OpenAI, Anthropic, Google), and universal across programming languages (Python, C, Rust, Go). Moreover, FPAs remain effective even when models are explicitly warned about the attack via robust system prompts. Finally, we explore positive, defensive uses of FPAs and discuss their broader implications for the reliability and safety of code-oriented LLMs.
title Trust Me, I Know This Function: Hijacking LLM Static Analysis using Bias
topic Machine Learning
Cryptography and Security
url https://arxiv.org/abs/2508.17361