PSM: Prompt Sensitivity Minimization via LLM-Guided Black-Box Optimization
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Jawad, Huseein, Brunel, Nicolas |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
par: Zhang, Chiyu, et autres
Publié: (2025)
par: Zhang, Chiyu, et autres
Publié: (2025)
TrailBlazer: History-Guided Reinforcement Learning for Black-Box LLM Jailbreaking
par: Yoon, Sung-Hoon, et autres
Publié: (2026)
par: Yoon, Sung-Hoon, et autres
Publié: (2026)
FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks
par: Chen, Bocheng, et autres
Publié: (2024)
par: Chen, Bocheng, et autres
Publié: (2024)
Rethinking LLM Watermark Detection in Black-Box Settings: A Non-Intrusive Third-Party Framework
par: Wang, Zhuoshang, et autres
Publié: (2026)
par: Wang, Zhuoshang, et autres
Publié: (2026)
Prompt Optimization and Evaluation for LLM Automated Red Teaming
par: Freenor, Michael, et autres
Publié: (2025)
par: Freenor, Michael, et autres
Publié: (2025)
Black-Box Guardrail Reverse-engineering Attack
par: Yao, Hongwei, et autres
Publié: (2025)
par: Yao, Hongwei, et autres
Publié: (2025)
BinarySelect to Improve Accessibility of Black-Box Attack Research
par: Ghosh, Shatarupa, et autres
Publié: (2024)
par: Ghosh, Shatarupa, et autres
Publié: (2024)
RTD-Guard: A Black-Box Textual Adversarial Detection Framework via Replacement Token Detection
par: Zhu, He, et autres
Publié: (2026)
par: Zhu, He, et autres
Publié: (2026)
BiAxisAudit: A Novel Framework to Evaluate LLM Bias Across Prompt Sensitivity and Response-Layer Divergence
par: Gan, Jialing, et autres
Publié: (2026)
par: Gan, Jialing, et autres
Publié: (2026)
Cross-Lingual Summarization as a Black-Box Watermark Removal Attack
par: Ganesan, Gokul
Publié: (2025)
par: Ganesan, Gokul
Publié: (2025)
Auditing Black-Box LLM APIs with a Rank-Based Uniformity Test
par: Zhu, Xiaoyuan, et autres
Publié: (2025)
par: Zhu, Xiaoyuan, et autres
Publié: (2025)
SecureForge: Finding and Preventing Vulnerabilities in LLM-Generated Code via Prompt Optimization
par: Liu, Houjun, et autres
Publié: (2026)
par: Liu, Houjun, et autres
Publié: (2026)
EvoDefense: Co-Evolving Black-Box Defense with Large Language Models
par: Li, Yu, et autres
Publié: (2026)
par: Li, Yu, et autres
Publié: (2026)
AlienLM: Alienization of Language for API-Boundary Privacy in Black-Box LLMs
par: Kim, Jaehee, et autres
Publié: (2026)
par: Kim, Jaehee, et autres
Publié: (2026)
"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks
par: Wang, Libo
Publié: (2024)
par: Wang, Libo
Publié: (2024)
TRAP: Targeted Random Adversarial Prompt Honeypot for Black-Box Identification
par: Gubri, Martin, et autres
Publié: (2024)
par: Gubri, Martin, et autres
Publié: (2024)
SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
par: Li, Yucheng, et autres
Publié: (2025)
par: Li, Yucheng, et autres
Publié: (2025)
Robust LLM Watermarking with Minimal Semantic Distortion for IP Protection
par: Dang, Kieu, et autres
Publié: (2026)
par: Dang, Kieu, et autres
Publié: (2026)
Confidential Prompting: Privacy-preserving LLM Inference on Cloud
par: Li, Caihua, et autres
Publié: (2024)
par: Li, Caihua, et autres
Publié: (2024)
Raccoon: Prompt Extraction Benchmark of LLM-Integrated Applications
par: Wang, Junlin, et autres
Publié: (2024)
par: Wang, Junlin, et autres
Publié: (2024)
A Watermark for Black-Box Language Models
par: Bahri, Dara, et autres
Publié: (2024)
par: Bahri, Dara, et autres
Publié: (2024)
PAL: Proxy-Guided Black-Box Attack on Large Language Models
par: Sitawarin, Chawin, et autres
Publié: (2024)
par: Sitawarin, Chawin, et autres
Publié: (2024)
The Landscape of Prompt Injection Threats in LLM Agents: From Taxonomy to Analysis
par: Wang, Peiran, et autres
Publié: (2026)
par: Wang, Peiran, et autres
Publié: (2026)
Fingerprinting LLMs via Prompt Injection
par: Hu, Yuepeng, et autres
Publié: (2025)
par: Hu, Yuepeng, et autres
Publié: (2025)
Fun-tuning: Characterizing the Vulnerability of Proprietary LLMs to Optimization-based Prompt Injection Attacks via the Fine-Tuning Interface
par: Labunets, Andrey, et autres
Publié: (2025)
par: Labunets, Andrey, et autres
Publié: (2025)
Graph of Attacks: Improved Black-Box and Interpretable Jailbreaks for LLMs
par: Akbar-Tajari, Mohammad, et autres
Publié: (2025)
par: Akbar-Tajari, Mohammad, et autres
Publié: (2025)
Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections
par: Maloyan, Narek, et autres
Publié: (2025)
par: Maloyan, Narek, et autres
Publié: (2025)
SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts
par: Xin, Yuan, et autres
Publié: (2026)
par: Xin, Yuan, et autres
Publié: (2026)
Auto-Tuning Safety Guardrails for Black-Box Large Language Models
par: Abdulkadir, Perry
Publié: (2025)
par: Abdulkadir, Perry
Publié: (2025)
Effective and Efficient Jailbreaks of Black-Box LLMs with Cross-Behavior Attacks
par: Gohil, Vasudev
Publié: (2025)
par: Gohil, Vasudev
Publié: (2025)
Do Prompts Guarantee Safety? Mitigating Toxicity from LLM Generations through Subspace Intervention
par: Singh, Himanshu, et autres
Publié: (2026)
par: Singh, Himanshu, et autres
Publié: (2026)
SecureLLM: Using Compositionality to Build Provably Secure Language Models for Private, Sensitive, and Secret Data
par: Alabdulkareem, Abdulrahman, et autres
Publié: (2024)
par: Alabdulkareem, Abdulrahman, et autres
Publié: (2024)
Minimal Prompt Perturbations Lead to Code Vulnerabilities: Prompt Fragility and Hidden-State Signals in Coding LLMs
par: Sternfeld, Alexander, et autres
Publié: (2026)
par: Sternfeld, Alexander, et autres
Publié: (2026)
GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods
par: Huang, Ruixuan, et autres
Publié: (2025)
par: Huang, Ruixuan, et autres
Publié: (2025)
Graph of Attacks with Pruning: Optimizing Stealthy Jailbreak Prompt Generation for Enhanced LLM Content Moderation
par: Schwartz, Daniel, et autres
Publié: (2025)
par: Schwartz, Daniel, et autres
Publié: (2025)
Black-Box Opinion Manipulation Attacks to Retrieval-Augmented Generation of Large Language Models
par: Chen, Zhuo, et autres
Publié: (2024)
par: Chen, Zhuo, et autres
Publié: (2024)
Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs
par: Li, Xiang, et autres
Publié: (2025)
par: Li, Xiang, et autres
Publié: (2025)
Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation
par: Roh, Jaechul, et autres
Publié: (2025)
par: Roh, Jaechul, et autres
Publié: (2025)
GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient Analysis
par: Xie, Yueqi, et autres
Publié: (2024)
par: Xie, Yueqi, et autres
Publié: (2024)
PRSA: Prompt Stealing Attacks against Real-World Prompt Services
par: Yang, Yong, et autres
Publié: (2024)
par: Yang, Yong, et autres
Publié: (2024)
Documents similaires
-
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
par: Zhang, Chiyu, et autres
Publié: (2025) -
TrailBlazer: History-Guided Reinforcement Learning for Black-Box LLM Jailbreaking
par: Yoon, Sung-Hoon, et autres
Publié: (2026) -
FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks
par: Chen, Bocheng, et autres
Publié: (2024) -
Rethinking LLM Watermark Detection in Black-Box Settings: A Non-Intrusive Third-Party Framework
par: Wang, Zhuoshang, et autres
Publié: (2026) -
Prompt Optimization and Evaluation for LLM Automated Red Teaming
par: Freenor, Michael, et autres
Publié: (2025)