Salvato in:
| Autori principali: | Yao, Hongwei, Xia, Yun, Shao, Shuo, Shi, Haoran, Qiao, Tong, Wang, Cong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2511.04215 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks
di: Wang, Libo
Pubblicazione: (2024)
di: Wang, Libo
Pubblicazione: (2024)
Auto-Tuning Safety Guardrails for Black-Box Large Language Models
di: Abdulkadir, Perry
Pubblicazione: (2025)
di: Abdulkadir, Perry
Pubblicazione: (2025)
BinarySelect to Improve Accessibility of Black-Box Attack Research
di: Ghosh, Shatarupa, et al.
Pubblicazione: (2024)
di: Ghosh, Shatarupa, et al.
Pubblicazione: (2024)
Reading Between the Lines: Towards Reliable Black-box LLM Fingerprinting via Zeroth-order Gradient Estimation
di: Shao, Shuo, et al.
Pubblicazione: (2025)
di: Shao, Shuo, et al.
Pubblicazione: (2025)
Cross-Lingual Summarization as a Black-Box Watermark Removal Attack
di: Ganesan, Gokul
Pubblicazione: (2025)
di: Ganesan, Gokul
Pubblicazione: (2025)
OpenGuardrails: A Configurable, Unified, and Scalable Guardrails Platform for Large Language Models
di: Wang, Thomas, et al.
Pubblicazione: (2025)
di: Wang, Thomas, et al.
Pubblicazione: (2025)
Bag of Tricks for Subverting Reasoning-based Safety Guardrails
di: Chen, Shuo, et al.
Pubblicazione: (2025)
di: Chen, Shuo, et al.
Pubblicazione: (2025)
TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts
di: Chu, Hua-Rong, et al.
Pubblicazione: (2026)
di: Chu, Hua-Rong, et al.
Pubblicazione: (2026)
Triaging Threats to Specialized Guardrails
di: Mo, Wenjie Jacky, et al.
Pubblicazione: (2026)
di: Mo, Wenjie Jacky, et al.
Pubblicazione: (2026)
Graph of Attacks: Improved Black-Box and Interpretable Jailbreaks for LLMs
di: Akbar-Tajari, Mohammad, et al.
Pubblicazione: (2025)
di: Akbar-Tajari, Mohammad, et al.
Pubblicazione: (2025)
FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks
di: Chen, Bocheng, et al.
Pubblicazione: (2024)
di: Chen, Bocheng, et al.
Pubblicazione: (2024)
Effective and Efficient Jailbreaks of Black-Box LLMs with Cross-Behavior Attacks
di: Gohil, Vasudev
Pubblicazione: (2025)
di: Gohil, Vasudev
Pubblicazione: (2025)
Eguard: Defending LLM Embeddings Against Inversion Attacks via Text Mutual Information Optimization
di: Liu, Tiantian, et al.
Pubblicazione: (2024)
di: Liu, Tiantian, et al.
Pubblicazione: (2024)
Interpretable LLM Guardrails via Sparse Representation Steering
di: He, Zeqing, et al.
Pubblicazione: (2025)
di: He, Zeqing, et al.
Pubblicazione: (2025)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
di: Mehrotra, Anay, et al.
Pubblicazione: (2023)
di: Mehrotra, Anay, et al.
Pubblicazione: (2023)
Black-Box Opinion Manipulation Attacks to Retrieval-Augmented Generation of Large Language Models
di: Chen, Zhuo, et al.
Pubblicazione: (2024)
di: Chen, Zhuo, et al.
Pubblicazione: (2024)
SoK: Large Language Model Copyright Auditing via Fingerprinting
di: Shao, Shuo, et al.
Pubblicazione: (2025)
di: Shao, Shuo, et al.
Pubblicazione: (2025)
SGuard-v1: Safety Guardrail for Large Language Models
di: Lee, JoonHo, et al.
Pubblicazione: (2025)
di: Lee, JoonHo, et al.
Pubblicazione: (2025)
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
di: Zhang, Chiyu, et al.
Pubblicazione: (2025)
di: Zhang, Chiyu, et al.
Pubblicazione: (2025)
PAL: Proxy-Guided Black-Box Attack on Large Language Models
di: Sitawarin, Chawin, et al.
Pubblicazione: (2024)
di: Sitawarin, Chawin, et al.
Pubblicazione: (2024)
PSM: Prompt Sensitivity Minimization via LLM-Guided Black-Box Optimization
di: Jawad, Huseein, et al.
Pubblicazione: (2025)
di: Jawad, Huseein, et al.
Pubblicazione: (2025)
EvoDefense: Co-Evolving Black-Box Defense with Large Language Models
di: Li, Yu, et al.
Pubblicazione: (2026)
di: Li, Yu, et al.
Pubblicazione: (2026)
AlienLM: Alienization of Language for API-Boundary Privacy in Black-Box LLMs
di: Kim, Jaehee, et al.
Pubblicazione: (2026)
di: Kim, Jaehee, et al.
Pubblicazione: (2026)
MPMA: Preference Manipulation Attack Against Model Context Protocol
di: Wang, Zihan, et al.
Pubblicazione: (2025)
di: Wang, Zihan, et al.
Pubblicazione: (2025)
Rethinking LLM Watermark Detection in Black-Box Settings: A Non-Intrusive Third-Party Framework
di: Wang, Zhuoshang, et al.
Pubblicazione: (2026)
di: Wang, Zhuoshang, et al.
Pubblicazione: (2026)
Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
di: Wang, Haoran, et al.
Pubblicazione: (2023)
di: Wang, Haoran, et al.
Pubblicazione: (2023)
BadLingual: A Novel Lingual-Backdoor Attack against Large Language Models
di: Wang, Zihan, et al.
Pubblicazione: (2025)
di: Wang, Zihan, et al.
Pubblicazione: (2025)
A Watermark for Black-Box Language Models
di: Bahri, Dara, et al.
Pubblicazione: (2024)
di: Bahri, Dara, et al.
Pubblicazione: (2024)
OneShield -- the Next Generation of LLM Guardrails
di: DeLuca, Chad, et al.
Pubblicazione: (2025)
di: DeLuca, Chad, et al.
Pubblicazione: (2025)
ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models
di: Zhao, Yunhan, et al.
Pubblicazione: (2026)
di: Zhao, Yunhan, et al.
Pubblicazione: (2026)
Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation
di: Huang, Tiansheng, et al.
Pubblicazione: (2025)
di: Huang, Tiansheng, et al.
Pubblicazione: (2025)
Q-FAKER: Query-free Hard Black-box Attack via Controlled Generation
di: Na, CheolWon, et al.
Pubblicazione: (2025)
di: Na, CheolWon, et al.
Pubblicazione: (2025)
RTD-Guard: A Black-Box Textual Adversarial Detection Framework via Replacement Token Detection
di: Zhu, He, et al.
Pubblicazione: (2026)
di: Zhu, He, et al.
Pubblicazione: (2026)
TRAP: Targeted Random Adversarial Prompt Honeypot for Black-Box Identification
di: Gubri, Martin, et al.
Pubblicazione: (2024)
di: Gubri, Martin, et al.
Pubblicazione: (2024)
Proof-of-Guardrail in AI Agents and What (Not) to Trust from It
di: Jin, Xisen, et al.
Pubblicazione: (2026)
di: Jin, Xisen, et al.
Pubblicazione: (2026)
NeuroFilter: Privacy Guardrails for Conversational LLM Agents
di: Das, Saswat, et al.
Pubblicazione: (2026)
di: Das, Saswat, et al.
Pubblicazione: (2026)
Reversible Jump Attack to Textual Classifiers with Modification Reduction
di: Ni, Mingze, et al.
Pubblicazione: (2024)
di: Ni, Mingze, et al.
Pubblicazione: (2024)
$PD^3F$: A Pluggable and Dynamic DoS-Defense Framework Against Resource Consumption Attacks Targeting Large Language Models
di: Zhang, Yuanhe, et al.
Pubblicazione: (2025)
di: Zhang, Yuanhe, et al.
Pubblicazione: (2025)
TrailBlazer: History-Guided Reinforcement Learning for Black-Box LLM Jailbreaking
di: Yoon, Sung-Hoon, et al.
Pubblicazione: (2026)
di: Yoon, Sung-Hoon, et al.
Pubblicazione: (2026)
Privacy in Large Language Models: Attacks, Defenses and Future Directions
di: Li, Haoran, et al.
Pubblicazione: (2023)
di: Li, Haoran, et al.
Pubblicazione: (2023)
Documenti analoghi
-
"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks
di: Wang, Libo
Pubblicazione: (2024) -
Auto-Tuning Safety Guardrails for Black-Box Large Language Models
di: Abdulkadir, Perry
Pubblicazione: (2025) -
BinarySelect to Improve Accessibility of Black-Box Attack Research
di: Ghosh, Shatarupa, et al.
Pubblicazione: (2024) -
Reading Between the Lines: Towards Reliable Black-box LLM Fingerprinting via Zeroth-order Gradient Estimation
di: Shao, Shuo, et al.
Pubblicazione: (2025) -
Cross-Lingual Summarization as a Black-Box Watermark Removal Attack
di: Ganesan, Gokul
Pubblicazione: (2025)