Salvato in:
| Autori principali: | Kim, Jinhwa, Harris, Ian G. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2508.10031 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning
di: Wang, Yanbo, et al.
Pubblicazione: (2026)
di: Wang, Yanbo, et al.
Pubblicazione: (2026)
Medical Malice: A Dataset for Context-Aware Safety in Healthcare LLMs
di: D'addario, Andrew Maranhão Ventura
Pubblicazione: (2025)
di: D'addario, Andrew Maranhão Ventura
Pubblicazione: (2025)
PISanitizer: Preventing Prompt Injection to Long-Context LLMs via Prompt Sanitization
di: Geng, Runpeng, et al.
Pubblicazione: (2025)
di: Geng, Runpeng, et al.
Pubblicazione: (2025)
PrimeGuard: Safe and Helpful LLMs through Tuning-Free Routing
di: Manczak, Blazej, et al.
Pubblicazione: (2024)
di: Manczak, Blazej, et al.
Pubblicazione: (2024)
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
di: Shao, Zedian, et al.
Pubblicazione: (2024)
di: Shao, Zedian, et al.
Pubblicazione: (2024)
Defend LLMs Through Self-Consciousness
di: Huang, Boshi, et al.
Pubblicazione: (2025)
di: Huang, Boshi, et al.
Pubblicazione: (2025)
Conversational Context Classification: A Representation Engineering Approach
di: Pan, Jonathan
Pubblicazione: (2026)
di: Pan, Jonathan
Pubblicazione: (2026)
Protecting Your LLMs with Information Bottleneck
di: Liu, Zichuan, et al.
Pubblicazione: (2024)
di: Liu, Zichuan, et al.
Pubblicazione: (2024)
SciSafeEval: A Comprehensive Benchmark for Safety Alignment of Large Language Models in Scientific Tasks
di: Li, Tianhao, et al.
Pubblicazione: (2024)
di: Li, Tianhao, et al.
Pubblicazione: (2024)
Universal and Context-Independent Triggers for Precise Control of LLM Outputs
di: Liang, Jiashuo, et al.
Pubblicazione: (2024)
di: Liang, Jiashuo, et al.
Pubblicazione: (2024)
Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs
di: Xu, Zhao, et al.
Pubblicazione: (2024)
di: Xu, Zhao, et al.
Pubblicazione: (2024)
The Ethics of Interaction: Mitigating Security Threats in LLMs
di: Kumar, Ashutosh, et al.
Pubblicazione: (2024)
di: Kumar, Ashutosh, et al.
Pubblicazione: (2024)
CCJA: Context-Coherent Jailbreak Attack for Aligned Large Language Models
di: Zhou, Guanghao, et al.
Pubblicazione: (2025)
di: Zhou, Guanghao, et al.
Pubblicazione: (2025)
Strategic Deflection: Defending LLMs from Logit Manipulation
di: Rachidy, Yassine, et al.
Pubblicazione: (2025)
di: Rachidy, Yassine, et al.
Pubblicazione: (2025)
PARASITE: Conditional System Prompt Poisoning to Hijack LLMs
di: Pham, Viet, et al.
Pubblicazione: (2025)
di: Pham, Viet, et al.
Pubblicazione: (2025)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
di: Liu, Fan, et al.
Pubblicazione: (2024)
di: Liu, Fan, et al.
Pubblicazione: (2024)
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
di: Mireshghallah, Niloofar, et al.
Pubblicazione: (2023)
di: Mireshghallah, Niloofar, et al.
Pubblicazione: (2023)
In-Context Representation Hijacking
di: Yona, Itay, et al.
Pubblicazione: (2025)
di: Yona, Itay, et al.
Pubblicazione: (2025)
Graph of Attacks: Improved Black-Box and Interpretable Jailbreaks for LLMs
di: Akbar-Tajari, Mohammad, et al.
Pubblicazione: (2025)
di: Akbar-Tajari, Mohammad, et al.
Pubblicazione: (2025)
bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs
di: Ji, Wence, et al.
Pubblicazione: (2025)
di: Ji, Wence, et al.
Pubblicazione: (2025)
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
di: Liu, Songyang, et al.
Pubblicazione: (2025)
di: Liu, Songyang, et al.
Pubblicazione: (2025)
Topology Matters: Measuring Memory Leakage in Multi-Agent LLMs
di: Liu, Jinbo, et al.
Pubblicazione: (2025)
di: Liu, Jinbo, et al.
Pubblicazione: (2025)
Feint and Attack: Attention-Based Strategies for Jailbreaking and Protecting LLMs
di: Pu, Rui, et al.
Pubblicazione: (2024)
di: Pu, Rui, et al.
Pubblicazione: (2024)
SeqAR: Jailbreak LLMs with Sequential Auto-Generated Characters
di: Yang, Yan, et al.
Pubblicazione: (2024)
di: Yang, Yan, et al.
Pubblicazione: (2024)
Waterfall: Framework for Robust and Scalable Text Watermarking and Provenance for LLMs
di: Lau, Gregory Kang Ruey, et al.
Pubblicazione: (2024)
di: Lau, Gregory Kang Ruey, et al.
Pubblicazione: (2024)
Sandwich attack: Multi-language Mixture Adaptive Attack on LLMs
di: Upadhayay, Bibek, et al.
Pubblicazione: (2024)
di: Upadhayay, Bibek, et al.
Pubblicazione: (2024)
Less is More: Sparse Watermarking in LLMs with Enhanced Text Quality
di: Hoang, Duy C., et al.
Pubblicazione: (2024)
di: Hoang, Duy C., et al.
Pubblicazione: (2024)
Rapid Optimization for Jailbreaking LLMs via Subconscious Exploitation and Echopraxia
di: Shen, Guangyu, et al.
Pubblicazione: (2024)
di: Shen, Guangyu, et al.
Pubblicazione: (2024)
Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation
di: Zhao, Shuai, et al.
Pubblicazione: (2024)
di: Zhao, Shuai, et al.
Pubblicazione: (2024)
Swiss-Bench 003: Evaluating LLM Reliability and Adversarial Security for Swiss Regulatory Contexts
di: Uenal, Fatih
Pubblicazione: (2026)
di: Uenal, Fatih
Pubblicazione: (2026)
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs
di: Hu, Xiaomeng, et al.
Pubblicazione: (2025)
di: Hu, Xiaomeng, et al.
Pubblicazione: (2025)
ShieldLearner: A New Paradigm for Jailbreak Attack Defense in LLMs
di: Ni, Ziyi, et al.
Pubblicazione: (2025)
di: Ni, Ziyi, et al.
Pubblicazione: (2025)
Effective and Efficient Jailbreaks of Black-Box LLMs with Cross-Behavior Attacks
di: Gohil, Vasudev
Pubblicazione: (2025)
di: Gohil, Vasudev
Pubblicazione: (2025)
Dagger Behind Smile: Fool LLMs with a Happy Ending Story
di: Song, Xurui, et al.
Pubblicazione: (2025)
di: Song, Xurui, et al.
Pubblicazione: (2025)
What Features in Prompts Jailbreak LLMs? Investigating the Mechanisms Behind Attacks
di: Kirch, Nathalie, et al.
Pubblicazione: (2024)
di: Kirch, Nathalie, et al.
Pubblicazione: (2024)
Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers
di: Wei, Jiali, et al.
Pubblicazione: (2026)
di: Wei, Jiali, et al.
Pubblicazione: (2026)
RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent
di: Xu, Huiyu, et al.
Pubblicazione: (2024)
di: Xu, Huiyu, et al.
Pubblicazione: (2024)
Injecting Falsehoods: Adversarial Man-in-the-Middle Attacks Undermining Factual Recall in LLMs
di: Fastowski, Alina, et al.
Pubblicazione: (2025)
di: Fastowski, Alina, et al.
Pubblicazione: (2025)
LoopLLM: Transferable Energy-Latency Attacks in LLMs via Repetitive Generation
di: Li, Xingyu, et al.
Pubblicazione: (2025)
di: Li, Xingyu, et al.
Pubblicazione: (2025)
DuFFin: A Dual-Level Fingerprinting Framework for LLMs IP Protection
di: Yan, Yuliang, et al.
Pubblicazione: (2025)
di: Yan, Yuliang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning
di: Wang, Yanbo, et al.
Pubblicazione: (2026) -
Medical Malice: A Dataset for Context-Aware Safety in Healthcare LLMs
di: D'addario, Andrew Maranhão Ventura
Pubblicazione: (2025) -
PISanitizer: Preventing Prompt Injection to Long-Context LLMs via Prompt Sanitization
di: Geng, Runpeng, et al.
Pubblicazione: (2025) -
PrimeGuard: Safe and Helpful LLMs through Tuning-Free Routing
di: Manczak, Blazej, et al.
Pubblicazione: (2024) -
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
di: Shao, Zedian, et al.
Pubblicazione: (2024)