Gespeichert in:
| 1. Verfasser: | Moss, Robert J. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2408.08899 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
von: Mehrotra, Anay, et al.
Veröffentlicht: (2023)
von: Mehrotra, Anay, et al.
Veröffentlicht: (2023)
PAL: Proxy-Guided Black-Box Attack on Large Language Models
von: Sitawarin, Chawin, et al.
Veröffentlicht: (2024)
von: Sitawarin, Chawin, et al.
Veröffentlicht: (2024)
TRAP: Targeted Random Adversarial Prompt Honeypot for Black-Box Identification
von: Gubri, Martin, et al.
Veröffentlicht: (2024)
von: Gubri, Martin, et al.
Veröffentlicht: (2024)
Exploiting Class Probabilities for Black-box Sentence-level Attacks
von: Moraffah, Raha, et al.
Veröffentlicht: (2024)
von: Moraffah, Raha, et al.
Veröffentlicht: (2024)
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations
von: Liang, Buyun, et al.
Veröffentlicht: (2025)
von: Liang, Buyun, et al.
Veröffentlicht: (2025)
REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations
von: Liang, Buyun, et al.
Veröffentlicht: (2026)
von: Liang, Buyun, et al.
Veröffentlicht: (2026)
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
von: Cao, Bochuan, et al.
Veröffentlicht: (2023)
von: Cao, Bochuan, et al.
Veröffentlicht: (2023)
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening
von: Zhang, Mohan, et al.
Veröffentlicht: (2026)
von: Zhang, Mohan, et al.
Veröffentlicht: (2026)
Exposing LLM Safety Gaps Through Mathematical Encoding:New Attacks and Systematic Analysis
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization
von: Tang, Haochun, et al.
Veröffentlicht: (2026)
von: Tang, Haochun, et al.
Veröffentlicht: (2026)
BlackDAN: A Black-Box Multi-Objective Approach for Effective and Contextual Jailbreaking of Large Language Models
von: Wang, Xinyuan, et al.
Veröffentlicht: (2024)
von: Wang, Xinyuan, et al.
Veröffentlicht: (2024)
Graph of Attacks: Improved Black-Box and Interpretable Jailbreaks for LLMs
von: Akbar-Tajari, Mohammad, et al.
Veröffentlicht: (2025)
von: Akbar-Tajari, Mohammad, et al.
Veröffentlicht: (2025)
Effective and Efficient Jailbreaks of Black-Box LLMs with Cross-Behavior Attacks
von: Gohil, Vasudev
Veröffentlicht: (2025)
von: Gohil, Vasudev
Veröffentlicht: (2025)
Membership Inference Attacks on LLM-based Recommender Systems
von: He, Jiajie, et al.
Veröffentlicht: (2025)
von: He, Jiajie, et al.
Veröffentlicht: (2025)
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
von: Liu, Yupei, et al.
Veröffentlicht: (2023)
von: Liu, Yupei, et al.
Veröffentlicht: (2023)
PBa-LLM: Privacy- and Bias-aware NLP using Named-Entity Recognition (NER)
von: Mancera, Gonzalo, et al.
Veröffentlicht: (2025)
von: Mancera, Gonzalo, et al.
Veröffentlicht: (2025)
BinaryShield: Cross-Service Threat Intelligence in LLM Services using Privacy-Preserving Fingerprints
von: Gill, Waris, et al.
Veröffentlicht: (2025)
von: Gill, Waris, et al.
Veröffentlicht: (2025)
The Resurgence of GCG Adversarial Attacks on Large Language Models
von: Tan, Yuting, et al.
Veröffentlicht: (2025)
von: Tan, Yuting, et al.
Veröffentlicht: (2025)
SurvAttack: Black-Box Attack On Survival Models through Ontology-Informed EHR Perturbation
von: Kerdabadi, Mohsen Nayebi, et al.
Veröffentlicht: (2024)
von: Kerdabadi, Mohsen Nayebi, et al.
Veröffentlicht: (2024)
SoK: Pitfalls in Evaluating Black-Box Attacks
von: Suya, Fnu, et al.
Veröffentlicht: (2023)
von: Suya, Fnu, et al.
Veröffentlicht: (2023)
Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data
von: Baumgärtner, Tim, et al.
Veröffentlicht: (2024)
von: Baumgärtner, Tim, et al.
Veröffentlicht: (2024)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
von: Shao, Zedian, et al.
Veröffentlicht: (2024)
von: Shao, Zedian, et al.
Veröffentlicht: (2024)
HSF: Defending against Jailbreak Attacks with Hidden State Filtering
von: Qian, Cheng, et al.
Veröffentlicht: (2024)
von: Qian, Cheng, et al.
Veröffentlicht: (2024)
Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking
von: Fang, Zhicheng, et al.
Veröffentlicht: (2026)
von: Fang, Zhicheng, et al.
Veröffentlicht: (2026)
Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks
von: Struppek, Lukas, et al.
Veröffentlicht: (2026)
von: Struppek, Lukas, et al.
Veröffentlicht: (2026)
Bypassing the Safety Training of Open-Source LLMs with Priming Attacks
von: Vega, Jason, et al.
Veröffentlicht: (2023)
von: Vega, Jason, et al.
Veröffentlicht: (2023)
On Adversarial Robustness of Language Models in Transfer Learning
von: Turbal, Bohdan, et al.
Veröffentlicht: (2024)
von: Turbal, Bohdan, et al.
Veröffentlicht: (2024)
Towards Understanding the Fragility of Multilingual LLMs against Fine-Tuning Attacks
von: Poppi, Samuele, et al.
Veröffentlicht: (2024)
von: Poppi, Samuele, et al.
Veröffentlicht: (2024)
Adversarial Attacks on Parts of Speech: An Empirical Study in Text-to-Image Generation
von: Shahariar, G M, et al.
Veröffentlicht: (2024)
von: Shahariar, G M, et al.
Veröffentlicht: (2024)
On Evaluating The Performance of Watermarked Machine-Generated Texts Under Adversarial Attacks
von: Liu, Zesen, et al.
Veröffentlicht: (2024)
von: Liu, Zesen, et al.
Veröffentlicht: (2024)
EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage
von: Liao, Zeyi, et al.
Veröffentlicht: (2024)
von: Liao, Zeyi, et al.
Veröffentlicht: (2024)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
Revealing Weaknesses in Text Watermarking Through Self-Information Rewrite Attacks
von: Cheng, Yixin, et al.
Veröffentlicht: (2025)
von: Cheng, Yixin, et al.
Veröffentlicht: (2025)
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
von: Chen, Taiye, et al.
Veröffentlicht: (2025)
von: Chen, Taiye, et al.
Veröffentlicht: (2025)
Attack and defense techniques in large language models: A survey and new perspectives
von: Liao, Zhiyu, et al.
Veröffentlicht: (2025)
von: Liao, Zhiyu, et al.
Veröffentlicht: (2025)
MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation
von: Jiang, Weisen, et al.
Veröffentlicht: (2025)
von: Jiang, Weisen, et al.
Veröffentlicht: (2025)
Checkpoint-GCG: Auditing and Attacking Fine-Tuning-Based Prompt Injection Defenses
von: Yang, Xiaoxue, et al.
Veröffentlicht: (2025)
von: Yang, Xiaoxue, et al.
Veröffentlicht: (2025)
Black-Box Opinion Manipulation Attacks to Retrieval-Augmented Generation of Large Language Models
von: Chen, Zhuo, et al.
Veröffentlicht: (2024)
von: Chen, Zhuo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
von: Mehrotra, Anay, et al.
Veröffentlicht: (2023) -
PAL: Proxy-Guided Black-Box Attack on Large Language Models
von: Sitawarin, Chawin, et al.
Veröffentlicht: (2024) -
TRAP: Targeted Random Adversarial Prompt Honeypot for Black-Box Identification
von: Gubri, Martin, et al.
Veröffentlicht: (2024) -
Exploiting Class Probabilities for Black-box Sentence-level Attacks
von: Moraffah, Raha, et al.
Veröffentlicht: (2024) -
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
von: Wang, Yifei, et al.
Veröffentlicht: (2024)