Trust Me, I Know This Function: Hijacking LLM Static Analysis using Bias
Fuente:
arXiv
Saved in:
| Main Authors: | Bernstein, Shir, Beste, David, Ayzenshteyn, Daniel, Schonherr, Lea, Mirsky, Yisroel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Best Defense is a Good Offense: Countering LLM-Powered Cyberattacks
by: Ayzenshteyn, Daniel, et al.
Published: (2024)
by: Ayzenshteyn, Daniel, et al.
Published: (2024)
Efficient Model Extraction via Boundary Sampling
by: Dor, Maor Biton, et al.
Published: (2024)
by: Dor, Maor Biton, et al.
Published: (2024)
PEAS: A Strategy for Crafting Transferable Adversarial Examples
by: Avraham, Bar, et al.
Published: (2024)
by: Avraham, Bar, et al.
Published: (2024)
Counter-Samples: A Stateless Strategy to Neutralize Black Box Adversarial Attacks
by: Bokobza, Roey, et al.
Published: (2024)
by: Bokobza, Roey, et al.
Published: (2024)
Transpose Attack: Stealing Datasets with Bidirectional Training
by: Amit, Guy, et al.
Published: (2023)
by: Amit, Guy, et al.
Published: (2023)
What Was Your Prompt? A Remote Keylogging Attack on AI Assistants
by: Weiss, Roy, et al.
Published: (2024)
by: Weiss, Roy, et al.
Published: (2024)
GAVEL: Towards Rule-Based Safety Through Activation Monitoring
by: Rozenfeld, Shir, et al.
Published: (2026)
by: Rozenfeld, Shir, et al.
Published: (2026)
Transferability Ranking of Adversarial Examples
by: Levy, Mosh, et al.
Published: (2022)
by: Levy, Mosh, et al.
Published: (2022)
One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries
by: Zloczower, Itay, et al.
Published: (2026)
by: Zloczower, Itay, et al.
Published: (2026)
Prompt Obfuscation for Large Language Models
by: Pape, David, et al.
Published: (2024)
by: Pape, David, et al.
Published: (2024)
Whispers in the Machine: Confidentiality in Agentic Systems
by: Evertz, Jonathan, et al.
Published: (2024)
by: Evertz, Jonathan, et al.
Published: (2024)
Memory Backdoor Attacks on Neural Networks
by: Luzon, Eden, et al.
Published: (2024)
by: Luzon, Eden, et al.
Published: (2024)
FuncPoison: Poisoning Function Library to Hijack Multi-agent Autonomous Driving Systems
by: Long, Yuzhen, et al.
Published: (2025)
by: Long, Yuzhen, et al.
Published: (2025)
Model Hijacking Attack in Federated Learning
by: Li, Zheng, et al.
Published: (2024)
by: Li, Zheng, et al.
Published: (2024)
Vera Verto: Multimodal Hijacking Attack
by: Zhang, Minxing, et al.
Published: (2024)
by: Zhang, Minxing, et al.
Published: (2024)
Osmosis Distillation: Model Hijacking with the Fewest Samples
by: Shi, Yuchen, et al.
Published: (2026)
by: Shi, Yuchen, et al.
Published: (2026)
ProxyPrints: From Database Breach to Spoof, A Plug-and-Play Defense for Biometric Systems
by: Hacmon, Yaniv, et al.
Published: (2025)
by: Hacmon, Yaniv, et al.
Published: (2025)
Seed Hijacking of LLM Sampling and Quantum Random Number Defense
by: You, Ziyang, et al.
Published: (2026)
by: You, Ziyang, et al.
Published: (2026)
Know Me by My Pulse: Toward Practical Continuous Authentication on Wearable Devices via Wrist-Worn PPG
by: Shao, Wei, et al.
Published: (2025)
by: Shao, Wei, et al.
Published: (2025)
Beyond Crash: Hijacking Your Autonomous Vehicle for Fun and Profit
by: Sun, Qi, et al.
Published: (2026)
by: Sun, Qi, et al.
Published: (2026)
Moshi Moshi? A Model Selection Hijacking Adversarial Attack
by: Petrucci, Riccardo, et al.
Published: (2025)
by: Petrucci, Riccardo, et al.
Published: (2025)
AutoRAN: Automated Hijacking of Safety Reasoning in Large Reasoning Models
by: Liang, Jiacheng, et al.
Published: (2025)
by: Liang, Jiacheng, et al.
Published: (2025)
Know Thy Judge: On the Robustness Meta-Evaluation of LLM Safety Judges
by: Eiras, Francisco, et al.
Published: (2025)
by: Eiras, Francisco, et al.
Published: (2025)
MIP against Agent: Malicious Image Patches Hijacking Multimodal OS Agents
by: Aichberger, Lukas, et al.
Published: (2025)
by: Aichberger, Lukas, et al.
Published: (2025)
AIJack: Let's Hijack AI! Security and Privacy Risk Simulator for Machine Learning
by: Takahashi, Hideaki
Published: (2023)
by: Takahashi, Hideaki
Published: (2023)
No More, No Less: Task Alignment in Terminal Agents
by: Mavali, Sina, et al.
Published: (2026)
by: Mavali, Sina, et al.
Published: (2026)
SplitOut: Out-of-the-Box Training-Hijacking Detection in Split Learning via Outlier Detection
by: Erdogan, Ege, et al.
Published: (2023)
by: Erdogan, Ege, et al.
Published: (2023)
DiffusionHijack: Supply-Chain PRNG Backdoor Attack on Diffusion Models and Quantum Random Number Defense
by: You, Ziyang, et al.
Published: (2026)
by: You, Ziyang, et al.
Published: (2026)
Love, Lies, and Language Models: Investigating AI's Role in Romance-Baiting Scams
by: Gressel, Gilad, et al.
Published: (2025)
by: Gressel, Gilad, et al.
Published: (2025)
UDora: A Unified Red Teaming Framework against LLM Agents by Dynamically Hijacking Their Own Reasoning
by: Zhang, Jiawei, et al.
Published: (2025)
by: Zhang, Jiawei, et al.
Published: (2025)
CHAI: Command Hijacking against embodied AI
by: Burbano, Luis, et al.
Published: (2025)
by: Burbano, Luis, et al.
Published: (2025)
HexaCoder: Secure Code Generation via Oracle-Guided Synthetic Training Data
by: Hajipour, Hossein, et al.
Published: (2024)
by: Hajipour, Hossein, et al.
Published: (2024)
Hijack Vertical Federated Learning Models As One Party
by: Qiu, Pengyu, et al.
Published: (2022)
by: Qiu, Pengyu, et al.
Published: (2022)
Beyond the Wrapper: Identifying Artifact Reliance in Static Malware Classifiers using TRUSTEE
by: Mohammed, Riyazuddin, et al.
Published: (2026)
by: Mohammed, Riyazuddin, et al.
Published: (2026)
In-Context Representation Hijacking
by: Yona, Itay, et al.
Published: (2025)
by: Yona, Itay, et al.
Published: (2025)
Empirical Analysis of Large Vision-Language Models against Goal Hijacking via Visual Prompt Injection
by: Kimura, Subaru, et al.
Published: (2024)
by: Kimura, Subaru, et al.
Published: (2024)
Hijacking Large Language Models via Adversarial In-Context Learning
by: Zhou, Xiangyu, et al.
Published: (2023)
by: Zhou, Xiangyu, et al.
Published: (2023)
R+R: Revisiting Static Feature-Based Android Malware Detection using Machine Learning
by: Alam, Md Tanvirul, et al.
Published: (2024)
by: Alam, Md Tanvirul, et al.
Published: (2024)
DeepTrust: Multi-Step Classification through Dissimilar Adversarial Representations for Robust Android Malware Detection
by: Pulido-Cortázar, Daniel, et al.
Published: (2025)
by: Pulido-Cortázar, Daniel, et al.
Published: (2025)
Image Hijacks: Adversarial Images can Control Generative Models at Runtime
by: Bailey, Luke, et al.
Published: (2023)
by: Bailey, Luke, et al.
Published: (2023)
Similar Items
-
The Best Defense is a Good Offense: Countering LLM-Powered Cyberattacks
by: Ayzenshteyn, Daniel, et al.
Published: (2024) -
Efficient Model Extraction via Boundary Sampling
by: Dor, Maor Biton, et al.
Published: (2024) -
PEAS: A Strategy for Crafting Transferable Adversarial Examples
by: Avraham, Bar, et al.
Published: (2024) -
Counter-Samples: A Stateless Strategy to Neutralize Black Box Adversarial Attacks
by: Bokobza, Roey, et al.
Published: (2024) -
Transpose Attack: Stealing Datasets with Bidirectional Training
by: Amit, Guy, et al.
Published: (2023)