Well, that escalated quickly: The Single-Turn Crescendo Attack (STCA)
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Aqrawi, Alan, Abbasi, Arian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
An indicator for effectiveness of text-to-image guardrails utilizing the Single-Turn Crescendo Attack (STCA)
von: Kwartler, Ted, et al.
Veröffentlicht: (2024)
von: Kwartler, Ted, et al.
Veröffentlicht: (2024)
Good Parenting is all you need -- Multi-agentic LLM Hallucination Mitigation
von: Kwartler, Ted, et al.
Veröffentlicht: (2024)
von: Kwartler, Ted, et al.
Veröffentlicht: (2024)
Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack
von: Russinovich, Mark, et al.
Veröffentlicht: (2024)
von: Russinovich, Mark, et al.
Veröffentlicht: (2024)
Let the Bees Find the Weak Spots: A Path Planning Perspective on Multi-Turn Jailbreak Attacks against LLMs
von: Liu, Yize, et al.
Veröffentlicht: (2025)
von: Liu, Yize, et al.
Veröffentlicht: (2025)
Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks
von: Gibbs, Tom, et al.
Veröffentlicht: (2024)
von: Gibbs, Tom, et al.
Veröffentlicht: (2024)
Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLM
von: Yang, Xikang, et al.
Veröffentlicht: (2024)
von: Yang, Xikang, et al.
Veröffentlicht: (2024)
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
Siren: A Learning-Based Multi-Turn Attack Framework for Simulating Real-World Human Jailbreak Behaviors
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
Membership Inference Attacks Against In-Context Learning
von: Wen, Rui, et al.
Veröffentlicht: (2024)
von: Wen, Rui, et al.
Veröffentlicht: (2024)
Black-Box Guardrail Reverse-engineering Attack
von: Yao, Hongwei, et al.
Veröffentlicht: (2025)
von: Yao, Hongwei, et al.
Veröffentlicht: (2025)
Task-Agnostic Detector for Insertion-Based Backdoor Attacks
von: Lyu, Weimin, et al.
Veröffentlicht: (2024)
von: Lyu, Weimin, et al.
Veröffentlicht: (2024)
Security Attacks on LLM-based Code Completion Tools
von: Cheng, Wen, et al.
Veröffentlicht: (2024)
von: Cheng, Wen, et al.
Veröffentlicht: (2024)
Prompt Stealing Attacks Against Large Language Models
von: Sha, Zeyang, et al.
Veröffentlicht: (2024)
von: Sha, Zeyang, et al.
Veröffentlicht: (2024)
SBFA: Single Sneaky Bit Flip Attack to Break Large Language Models
von: Guo, Jingkai, et al.
Veröffentlicht: (2025)
von: Guo, Jingkai, et al.
Veröffentlicht: (2025)
Data Extraction Attacks in Retrieval-Augmented Generation via Backdoors
von: Peng, Yuefeng, et al.
Veröffentlicht: (2024)
von: Peng, Yuefeng, et al.
Veröffentlicht: (2024)
Denial-of-Service Poisoning Attacks against Large Language Models
von: Gao, Kuofeng, et al.
Veröffentlicht: (2024)
von: Gao, Kuofeng, et al.
Veröffentlicht: (2024)
BinarySelect to Improve Accessibility of Black-Box Attack Research
von: Ghosh, Shatarupa, et al.
Veröffentlicht: (2024)
von: Ghosh, Shatarupa, et al.
Veröffentlicht: (2024)
Privacy in Large Language Models: Attacks, Defenses and Future Directions
von: Li, Haoran, et al.
Veröffentlicht: (2023)
von: Li, Haoran, et al.
Veröffentlicht: (2023)
MPMA: Preference Manipulation Attack Against Model Context Protocol
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
von: Chen, Yunhao, et al.
Veröffentlicht: (2025)
von: Chen, Yunhao, et al.
Veröffentlicht: (2025)
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue
von: Shen, Xinjie, et al.
Veröffentlicht: (2026)
von: Shen, Xinjie, et al.
Veröffentlicht: (2026)
PRSA: Prompt Stealing Attacks against Real-World Prompt Services
von: Yang, Yong, et al.
Veröffentlicht: (2024)
von: Yang, Yong, et al.
Veröffentlicht: (2024)
Harnessing Task Overload for Scalable Jailbreak Attacks on Large Language Models
von: Dong, Yiting, et al.
Veröffentlicht: (2024)
von: Dong, Yiting, et al.
Veröffentlicht: (2024)
No Two Devils Alike: Unveiling Distinct Mechanisms of Fine-tuning Attacks
von: Leong, Chak Tou, et al.
Veröffentlicht: (2024)
von: Leong, Chak Tou, et al.
Veröffentlicht: (2024)
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
von: Shen, Guobin, et al.
Veröffentlicht: (2025)
von: Shen, Guobin, et al.
Veröffentlicht: (2025)
RAG Safety: Exploring Knowledge Poisoning Attacks to Retrieval-Augmented Generation
von: Zhao, Tianzhe, et al.
Veröffentlicht: (2025)
von: Zhao, Tianzhe, et al.
Veröffentlicht: (2025)
Counterfactual Evaluation for Blind Attack Detection in LLM-based Evaluation Systems
von: Liu, Lijia, et al.
Veröffentlicht: (2025)
von: Liu, Lijia, et al.
Veröffentlicht: (2025)
Cross-Lingual Summarization as a Black-Box Watermark Removal Attack
von: Ganesan, Gokul
Veröffentlicht: (2025)
von: Ganesan, Gokul
Veröffentlicht: (2025)
Defending LLM Watermarking Against Spoofing Attacks with Contrastive Representation Learning
von: An, Li, et al.
Veröffentlicht: (2025)
von: An, Li, et al.
Veröffentlicht: (2025)
Securing Genomic Data Against Inference Attacks in Federated Learning Environments
von: Pathade, Chetan, et al.
Veröffentlicht: (2025)
von: Pathade, Chetan, et al.
Veröffentlicht: (2025)
Talk is (Not) Cheap: A Taxonomy and Benchmark Coverage Audit for LLM Attacks
von: Iyer, Karthik Raghu, et al.
Veröffentlicht: (2026)
von: Iyer, Karthik Raghu, et al.
Veröffentlicht: (2026)
Large Language Models are Good Attackers: Efficient and Stealthy Textual Backdoor Attacks
von: Li, Ziqiang, et al.
Veröffentlicht: (2024)
von: Li, Ziqiang, et al.
Veröffentlicht: (2024)
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
TuBA: Cross-Lingual Transferability of Backdoor Attacks in LLMs with Instruction Tuning
von: He, Xuanli, et al.
Veröffentlicht: (2024)
von: He, Xuanli, et al.
Veröffentlicht: (2024)
Enhance Robustness of Language Models Against Variation Attack through Graph Integration
von: Xiong, Zi, et al.
Veröffentlicht: (2024)
von: Xiong, Zi, et al.
Veröffentlicht: (2024)
Breaking Down the Defenses: A Comparative Survey of Attacks on Large Language Models
von: Chowdhury, Arijit Ghosh, et al.
Veröffentlicht: (2024)
von: Chowdhury, Arijit Ghosh, et al.
Veröffentlicht: (2024)
PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-Rails
von: Mangaokar, Neal, et al.
Veröffentlicht: (2024)
von: Mangaokar, Neal, et al.
Veröffentlicht: (2024)
Fact2Fiction: Targeted Poisoning Attack to Agentic Fact-checking System
von: He, Haorui, et al.
Veröffentlicht: (2025)
von: He, Haorui, et al.
Veröffentlicht: (2025)
Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections
von: Maloyan, Narek, et al.
Veröffentlicht: (2025)
von: Maloyan, Narek, et al.
Veröffentlicht: (2025)
RedTWIZ: Diverse LLM Red Teaming via Adaptive Attack Planning
von: Horal, Artur, et al.
Veröffentlicht: (2025)
von: Horal, Artur, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
An indicator for effectiveness of text-to-image guardrails utilizing the Single-Turn Crescendo Attack (STCA)
von: Kwartler, Ted, et al.
Veröffentlicht: (2024) -
Good Parenting is all you need -- Multi-agentic LLM Hallucination Mitigation
von: Kwartler, Ted, et al.
Veröffentlicht: (2024) -
Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack
von: Russinovich, Mark, et al.
Veröffentlicht: (2024) -
Let the Bees Find the Weak Spots: A Path Planning Perspective on Multi-Turn Jailbreak Attacks against LLMs
von: Liu, Yize, et al.
Veröffentlicht: (2025) -
Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks
von: Gibbs, Tom, et al.
Veröffentlicht: (2024)