How Few-shot Demonstrations Affect Prompt-based Defenses Against LLM Jailbreak Attacks
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Yanshu, Yang, Shuaishuai, He, Jingjing, Yang, Tong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
On the Validity of Traditional Vulnerability Scoring Systems for Adversarial Attacks against LLMs
por: Bahar, Atmane Ayoub Mansour, et al.
Publicado: (2024)
por: Bahar, Atmane Ayoub Mansour, et al.
Publicado: (2024)
On-Device Generative AI for GDPR-Compliant Visual Monitoring: Natural Language Alerts from Local Object Detection
por: Schappacher-Tilp, Gudrun, et al.
Publicado: (2026)
por: Schappacher-Tilp, Gudrun, et al.
Publicado: (2026)
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
por: Dang, Kieu, et al.
Publicado: (2025)
por: Dang, Kieu, et al.
Publicado: (2025)
AVEC: Bootstrapping Privacy for Local LLMs
por: Gaikwad, Madhava
Publicado: (2025)
por: Gaikwad, Madhava
Publicado: (2025)
Can Watermarked LLMs be Identified by Users via Crafted Prompts?
por: Liu, Aiwei, et al.
Publicado: (2024)
por: Liu, Aiwei, et al.
Publicado: (2024)
MarkLLM: An Open-Source Toolkit for LLM Watermarking
por: Pan, Leyi, et al.
Publicado: (2024)
por: Pan, Leyi, et al.
Publicado: (2024)
Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs
por: Hill, Brennen, et al.
Publicado: (2025)
por: Hill, Brennen, et al.
Publicado: (2025)
LegalGuardian: A Privacy-Preserving Framework for Secure Integration of Large Language Models in Legal Practice
por: Demir, M. Mikail, et al.
Publicado: (2025)
por: Demir, M. Mikail, et al.
Publicado: (2025)
AlignDP: Hybrid Differential Privacy with Rarity-Aware Protection for LLMs
por: Gaikwad, Madhava
Publicado: (2025)
por: Gaikwad, Madhava
Publicado: (2025)
BreakFun: Jailbreaking LLMs via Schema Exploitation
por: Oskooei, Amirkia Rafiei, et al.
Publicado: (2025)
por: Oskooei, Amirkia Rafiei, et al.
Publicado: (2025)
A Semantic Invariant Robust Watermark for Large Language Models
por: Liu, Aiwei, et al.
Publicado: (2023)
por: Liu, Aiwei, et al.
Publicado: (2023)
Digital Forgetting in Large Language Models: A Survey of Unlearning Methods
por: Blanco-Justicia, Alberto, et al.
Publicado: (2024)
por: Blanco-Justicia, Alberto, et al.
Publicado: (2024)
Shortcuts Arising from Contrast: Effective and Covert Clean-Label Attacks in Prompt-Based Learning
por: Xie, Xiaopeng, et al.
Publicado: (2024)
por: Xie, Xiaopeng, et al.
Publicado: (2024)
Powerful Training-Free Membership Inference Against Autoregressive Language Models
por: Ilić, David, et al.
Publicado: (2026)
por: Ilić, David, et al.
Publicado: (2026)
MarkDiffusion: An Open-Source Toolkit for Generative Watermarking of Latent Diffusion Models
por: Pan, Leyi, et al.
Publicado: (2025)
por: Pan, Leyi, et al.
Publicado: (2025)
Exploiting Web Search Tools of AI Agents for Data Exfiltration
por: Rall, Dennis, et al.
Publicado: (2025)
por: Rall, Dennis, et al.
Publicado: (2025)
Tool Receipts, Not Zero-Knowledge Proofs: Practical Hallucination Detection for AI Agents
por: Basu, Abhinaba
Publicado: (2026)
por: Basu, Abhinaba
Publicado: (2026)
A Method for Quantifying Human Risk and a Blueprint for LLM Integration
por: Canale, Giuseppe
Publicado: (2025)
por: Canale, Giuseppe
Publicado: (2025)
How Worrying Are Privacy Attacks Against Machine Learning?
por: Domingo-Ferrer, Josep
Publicado: (2025)
por: Domingo-Ferrer, Josep
Publicado: (2025)
Accelerating Suffix Jailbreak attacks with Prefix-Shared KV-cache
por: Wang, Xinhai, et al.
Publicado: (2026)
por: Wang, Xinhai, et al.
Publicado: (2026)
JPS: Jailbreak Multimodal Large Language Models with Collaborative Visual Perturbation and Textual Steering
por: Chen, Renmiao, et al.
Publicado: (2025)
por: Chen, Renmiao, et al.
Publicado: (2025)
Tatemae: Detecting Alignment Faking via Tool Selection in LLMs
por: Leonesi, Matteo, et al.
Publicado: (2026)
por: Leonesi, Matteo, et al.
Publicado: (2026)
Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks
por: Merves, Tyler H., et al.
Publicado: (2026)
por: Merves, Tyler H., et al.
Publicado: (2026)
Scaling Patterns in Adversarial Alignment: Evidence from Multi-LLM Jailbreak Experiments
por: Nathanson, Samuel, et al.
Publicado: (2025)
por: Nathanson, Samuel, et al.
Publicado: (2025)
CAMP: Cumulative Agentic Masking and Pruning for Privacy Protection in Multi-Turn LLM Conversations
por: Panjwani, Aman
Publicado: (2026)
por: Panjwani, Aman
Publicado: (2026)
Harmful Terms and Where to Find Them: Measuring and Modeling Unfavorable Financial Terms and Conditions in Shopping Websites at Scale
por: Tsai, Elisa, et al.
Publicado: (2025)
por: Tsai, Elisa, et al.
Publicado: (2025)
ChatGPT Based Data Augmentation for Improved Parameter-Efficient Debiasing of LLMs
por: Han, Pengrui, et al.
Publicado: (2024)
por: Han, Pengrui, et al.
Publicado: (2024)
Distractor Injection Attacks on Large Reasoning Models: Characterization and Defense
por: Zhang, Zhehao, et al.
Publicado: (2025)
por: Zhang, Zhehao, et al.
Publicado: (2025)
Beyond the Benchmark: Innovative Defenses Against Prompt Injection Attacks
por: Shaheer, Safwan, et al.
Publicado: (2025)
por: Shaheer, Safwan, et al.
Publicado: (2025)
Toward Secure and Compliant AI: Organizational Standards and Protocols for NLP Model Lifecycle Management
por: Arora, Sunil, et al.
Publicado: (2025)
por: Arora, Sunil, et al.
Publicado: (2025)
BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders
por: DeLeeuw, Caleb
Publicado: (2026)
por: DeLeeuw, Caleb
Publicado: (2026)
KidsNanny: A Two-Stage Multimodal Content Moderation Pipeline Integrating Visual Classification, Object Detection, OCR, and Contextual Reasoning for Child Safety
por: Panchal, Viraj, et al.
Publicado: (2026)
por: Panchal, Viraj, et al.
Publicado: (2026)
Logits of API-Protected LLMs Leak Proprietary Information
por: Finlayson, Matthew, et al.
Publicado: (2024)
por: Finlayson, Matthew, et al.
Publicado: (2024)
On Adversarial Examples for Text Classification by Perturbing Latent Representations
por: Sooksatra, Korn, et al.
Publicado: (2024)
por: Sooksatra, Korn, et al.
Publicado: (2024)
Jailbreaking Attacks vs. Content Safety Filters: How Far Are We in the LLM Safety Arms Race?
por: Xin, Yuan, et al.
Publicado: (2025)
por: Xin, Yuan, et al.
Publicado: (2025)
Evaluating the Robustness of Large Language Model Safety Guardrails Against Adversarial Attacks
por: Young, Richard J.
Publicado: (2025)
por: Young, Richard J.
Publicado: (2025)
The Ethics Engine: A Modular Pipeline for Accessible Psychometric Assessment of Large Language Models
por: Van Clief, Jake, et al.
Publicado: (2025)
por: Van Clief, Jake, et al.
Publicado: (2025)
Watermarking for AI Content Detection: A Review on Text, Visual, and Audio Modalities
por: Cao, Lele
Publicado: (2025)
por: Cao, Lele
Publicado: (2025)
The AI Fiction Paradox
por: Elkins, Katherine
Publicado: (2026)
por: Elkins, Katherine
Publicado: (2026)
AI-Powered Citation Auditing: A Zero-Assumption Protocol for Systematic Reference Verification in Academic Research
por: van Rensburg, L. J. Janse
Publicado: (2025)
por: van Rensburg, L. J. Janse
Publicado: (2025)
Ejemplares similares
-
On the Validity of Traditional Vulnerability Scoring Systems for Adversarial Attacks against LLMs
por: Bahar, Atmane Ayoub Mansour, et al.
Publicado: (2024) -
On-Device Generative AI for GDPR-Compliant Visual Monitoring: Natural Language Alerts from Local Object Detection
por: Schappacher-Tilp, Gudrun, et al.
Publicado: (2026) -
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
por: Dang, Kieu, et al.
Publicado: (2025) -
AVEC: Bootstrapping Privacy for Local LLMs
por: Gaikwad, Madhava
Publicado: (2025) -
Can Watermarked LLMs be Identified by Users via Crafted Prompts?
por: Liu, Aiwei, et al.
Publicado: (2024)