Can LLMs be Fooled? Investigating Vulnerabilities in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Abdali, Sara, He, Jia, Barberan, CJ, Anarfi, Richard |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Securing Large Language Models: Threats, Vulnerabilities and Responsible Practices
von: Abdali, Sara, et al.
Veröffentlicht: (2024)
von: Abdali, Sara, et al.
Veröffentlicht: (2024)
Decoding the AI Pen: Techniques and Challenges in Detecting AI-Generated Text
von: Abdali, Sara, et al.
Veröffentlicht: (2024)
von: Abdali, Sara, et al.
Veröffentlicht: (2024)
Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs
von: Price, Sara, et al.
Veröffentlicht: (2024)
von: Price, Sara, et al.
Veröffentlicht: (2024)
Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks
von: Qraitem, Maan, et al.
Veröffentlicht: (2024)
von: Qraitem, Maan, et al.
Veröffentlicht: (2024)
Can LLMs Patch Security Issues?
von: Alrashedy, Kamel, et al.
Veröffentlicht: (2023)
von: Alrashedy, Kamel, et al.
Veröffentlicht: (2023)
How Vulnerable Are Edge LLMs?
von: Ding, Ao, et al.
Veröffentlicht: (2026)
von: Ding, Ao, et al.
Veröffentlicht: (2026)
When and How to Fool Explainable Models (and Humans) with Adversarial Examples
von: Vadillo, Jon, et al.
Veröffentlicht: (2021)
von: Vadillo, Jon, et al.
Veröffentlicht: (2021)
Towards Secure Intelligent O-RAN Architecture: Vulnerabilities, Threats and Promising Technical Solutions using LLMs
von: Motalleb, Mojdeh Karbalaee, et al.
Veröffentlicht: (2024)
von: Motalleb, Mojdeh Karbalaee, et al.
Veröffentlicht: (2024)
Multi-Trigger Poisoning Amplifies Backdoor Vulnerabilities in LLMs
von: Sivapiromrat, Sanhanat, et al.
Veröffentlicht: (2025)
von: Sivapiromrat, Sanhanat, et al.
Veröffentlicht: (2025)
LLMStinger: Jailbreaking LLMs using RL fine-tuned LLMs
von: Jha, Piyush, et al.
Veröffentlicht: (2024)
von: Jha, Piyush, et al.
Veröffentlicht: (2024)
Fooling SHAP with Output Shuffling Attacks
von: Yuan, Jun, et al.
Veröffentlicht: (2024)
von: Yuan, Jun, et al.
Veröffentlicht: (2024)
Revisiting DeepFool: generalization and improvement
von: Abdollahpoorrostam, Alireza, et al.
Veröffentlicht: (2023)
von: Abdollahpoorrostam, Alireza, et al.
Veröffentlicht: (2023)
On Benchmarking Code LLMs for Android Malware Analysis
von: He, Yiling, et al.
Veröffentlicht: (2025)
von: He, Yiling, et al.
Veröffentlicht: (2025)
Can LLMs Handle WebShell Detection? Overcoming Detection Challenges with Behavioral Function-Aware Framework
von: Han, Feijiang, et al.
Veröffentlicht: (2025)
von: Han, Feijiang, et al.
Veröffentlicht: (2025)
Playing the Fool: Jailbreaking LLMs and Multimodal LLMs with Out-of-Distribution Strategy
von: Jeong, Joonhyun, et al.
Veröffentlicht: (2025)
von: Jeong, Joonhyun, et al.
Veröffentlicht: (2025)
ChatBug: A Common Vulnerability of Aligned LLMs Induced by Chat Templates
von: Jiang, Fengqing, et al.
Veröffentlicht: (2024)
von: Jiang, Fengqing, et al.
Veröffentlicht: (2024)
Transparency Attacks: How Imperceptible Image Layers Can Fool AI Perception
von: McKee, Forrest, et al.
Veröffentlicht: (2024)
von: McKee, Forrest, et al.
Veröffentlicht: (2024)
Turn-Based Structural Triggers: Prompt-Free Backdoors in Multi-Turn LLMs
von: Lu, Yiyang, et al.
Veröffentlicht: (2026)
von: Lu, Yiyang, et al.
Veröffentlicht: (2026)
Excessive Reasoning Attack on Reasoning LLMs
von: Si, Wai Man, et al.
Veröffentlicht: (2025)
von: Si, Wai Man, et al.
Veröffentlicht: (2025)
Towards Watermarking of Open-Source LLMs
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2025)
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2025)
Entropy-Guided Attention for Private LLMs
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2025)
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2025)
From Lab to Reality: A Practical Evaluation of Deep Learning Models and LLMs for Vulnerability Detection
von: Lu, Chaomeng, et al.
Veröffentlicht: (2025)
von: Lu, Chaomeng, et al.
Veröffentlicht: (2025)
Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
von: Panfilov, Alexander, et al.
Veröffentlicht: (2025)
von: Panfilov, Alexander, et al.
Veröffentlicht: (2025)
Can We Infer Confidential Properties of Training Data from LLMs?
von: Huang, Pengrun, et al.
Veröffentlicht: (2025)
von: Huang, Pengrun, et al.
Veröffentlicht: (2025)
SCoPE: Evaluating LLMs for Software Vulnerability Detection
von: Gonçalves, José, et al.
Veröffentlicht: (2024)
von: Gonçalves, José, et al.
Veröffentlicht: (2024)
Efficient Adversarial Training in LLMs with Continuous Attacks
von: Xhonneux, Sophie, et al.
Veröffentlicht: (2024)
von: Xhonneux, Sophie, et al.
Veröffentlicht: (2024)
Instruction Backdoor Attacks Against Customized LLMs
von: Zhang, Rui, et al.
Veröffentlicht: (2024)
von: Zhang, Rui, et al.
Veröffentlicht: (2024)
Surgical Repair of Insecure Code Generation in LLMs
von: Sandoval, Gustavo, et al.
Veröffentlicht: (2026)
von: Sandoval, Gustavo, et al.
Veröffentlicht: (2026)
OverThink: Slowdown Attacks on Reasoning LLMs
von: Kumar, Abhinav, et al.
Veröffentlicht: (2025)
von: Kumar, Abhinav, et al.
Veröffentlicht: (2025)
IF-GUIDE: Influence Function-Guided Detoxification of LLMs
von: Coalson, Zachary, et al.
Veröffentlicht: (2025)
von: Coalson, Zachary, et al.
Veröffentlicht: (2025)
SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert Identification
von: Lai, Zhenglin, et al.
Veröffentlicht: (2025)
von: Lai, Zhenglin, et al.
Veröffentlicht: (2025)
Can Differentially Private Fine-tuning LLMs Protect Against Privacy Attacks?
von: Du, Hao, et al.
Veröffentlicht: (2025)
von: Du, Hao, et al.
Veröffentlicht: (2025)
BaxBench: Can LLMs Generate Correct and Secure Backends?
von: Vero, Mark, et al.
Veröffentlicht: (2025)
von: Vero, Mark, et al.
Veröffentlicht: (2025)
SLIP: Securing LLMs IP Using Weights Decomposition
von: Refael, Yehonathan, et al.
Veröffentlicht: (2024)
von: Refael, Yehonathan, et al.
Veröffentlicht: (2024)
Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents
von: Kumar, Priyanshu, et al.
Veröffentlicht: (2024)
von: Kumar, Priyanshu, et al.
Veröffentlicht: (2024)
Pr$εε$mpt: Sanitizing Sensitive Prompts for LLMs
von: Chowdhury, Amrita Roy, et al.
Veröffentlicht: (2025)
von: Chowdhury, Amrita Roy, et al.
Veröffentlicht: (2025)
Adversarial Suffix Filtering: a Defense Pipeline for LLMs
von: Khachaturov, David, et al.
Veröffentlicht: (2025)
von: Khachaturov, David, et al.
Veröffentlicht: (2025)
Towards Harnessing the Power of LLMs for ABAC Policy Mining
von: Babasaheb, More Aayush, et al.
Veröffentlicht: (2025)
von: Babasaheb, More Aayush, et al.
Veröffentlicht: (2025)
Trident: Improving Malware Detection with LLMs and Behavioral Features
von: Saul, Rebecca, et al.
Veröffentlicht: (2026)
von: Saul, Rebecca, et al.
Veröffentlicht: (2026)
Can LLMs get help from other LLMs without revealing private information?
von: Hartmann, Florian, et al.
Veröffentlicht: (2024)
von: Hartmann, Florian, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Securing Large Language Models: Threats, Vulnerabilities and Responsible Practices
von: Abdali, Sara, et al.
Veröffentlicht: (2024) -
Decoding the AI Pen: Techniques and Challenges in Detecting AI-Generated Text
von: Abdali, Sara, et al.
Veröffentlicht: (2024) -
Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs
von: Price, Sara, et al.
Veröffentlicht: (2024) -
Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks
von: Qraitem, Maan, et al.
Veröffentlicht: (2024) -
Can LLMs Patch Security Issues?
von: Alrashedy, Kamel, et al.
Veröffentlicht: (2023)