SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
Fuente:
arXiv
Salvato in:
| Autori principali: | Robey, Alexander, Wong, Eric, Hassani, Hamed, Pappas, George J. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Jailbreaking Black Box Large Language Models in Twenty Queries
di: Chao, Patrick, et al.
Pubblicazione: (2023)
di: Chao, Patrick, et al.
Pubblicazione: (2023)
Defending Large Language Models against Jailbreak Attacks via Semantic Smoothing
di: Ji, Jiabao, et al.
Pubblicazione: (2024)
di: Ji, Jiabao, et al.
Pubblicazione: (2024)
Jailbreaking LLM-Controlled Robots
di: Robey, Alexander, et al.
Pubblicazione: (2024)
di: Robey, Alexander, et al.
Pubblicazione: (2024)
Towards Realistic Guarantees: A Probabilistic Certificate for SmoothLLM
di: Kumarappan, Adarsh, et al.
Pubblicazione: (2025)
di: Kumarappan, Adarsh, et al.
Pubblicazione: (2025)
Adversarial Reasoning at Jailbreaking Time
di: Sabbaghi, Mahdi, et al.
Pubblicazione: (2025)
di: Sabbaghi, Mahdi, et al.
Pubblicazione: (2025)
Conformal Prediction with Learned Features
di: Kiyani, Shayan, et al.
Pubblicazione: (2024)
di: Kiyani, Shayan, et al.
Pubblicazione: (2024)
Length Optimization in Conformal Prediction
di: Kiyani, Shayan, et al.
Pubblicazione: (2024)
di: Kiyani, Shayan, et al.
Pubblicazione: (2024)
Conformal Prediction Beyond the Seen: A Missing Mass Perspective for Uncertainty Quantification in Generative Models
di: Noorani, Sima, et al.
Pubblicazione: (2025)
di: Noorani, Sima, et al.
Pubblicazione: (2025)
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
di: Chao, Patrick, et al.
Pubblicazione: (2024)
di: Chao, Patrick, et al.
Pubblicazione: (2024)
Benchmarking Misuse Mitigation Against Covert Adversaries
di: Brown, Davis, et al.
Pubblicazione: (2025)
di: Brown, Davis, et al.
Pubblicazione: (2025)
Safety Guardrails for LLM-Enabled Robots
di: Ravichandran, Zachary, et al.
Pubblicazione: (2025)
di: Ravichandran, Zachary, et al.
Pubblicazione: (2025)
Decision Theoretic Foundations for Conformal Prediction: Optimal Uncertainty Quantification for Risk-Averse Agents
di: Kiyani, Shayan, et al.
Pubblicazione: (2025)
di: Kiyani, Shayan, et al.
Pubblicazione: (2025)
When to Trust the Cheap Check: Weak and Strong Verification for Reasoning
di: Kiyani, Shayan, et al.
Pubblicazione: (2026)
di: Kiyani, Shayan, et al.
Pubblicazione: (2026)
Robust Decision Making with Partially Calibrated Forecasts
di: Kiyani, Shayan, et al.
Pubblicazione: (2025)
di: Kiyani, Shayan, et al.
Pubblicazione: (2025)
Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks
di: Zhou, Andy, et al.
Pubblicazione: (2024)
di: Zhou, Andy, et al.
Pubblicazione: (2024)
Temporal Difference Learning with Compressed Updates: Error-Feedback meets Reinforcement Learning
di: Mitra, Aritra, et al.
Pubblicazione: (2023)
di: Mitra, Aritra, et al.
Pubblicazione: (2023)
Conformal Inference under High-Dimensional Covariate Shifts via Likelihood-Ratio Regularization
di: Joshi, Sunay, et al.
Pubblicazione: (2025)
di: Joshi, Sunay, et al.
Pubblicazione: (2025)
Human-AI Collaborative Uncertainty Quantification
di: Noorani, Sima, et al.
Pubblicazione: (2025)
di: Noorani, Sima, et al.
Pubblicazione: (2025)
Adversarial Attacks on Robotic Vision Language Action Models
di: Jones, Eliot Krzysztof, et al.
Pubblicazione: (2025)
di: Jones, Eliot Krzysztof, et al.
Pubblicazione: (2025)
Evaluating the Performance of Large Language Models via Debates
di: Moniri, Behrad, et al.
Pubblicazione: (2024)
di: Moniri, Behrad, et al.
Pubblicazione: (2024)
Algorithms for Adversarially Robust Deep Learning
di: Robey, Alexander
Pubblicazione: (2025)
di: Robey, Alexander
Pubblicazione: (2025)
BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks
di: Anupam, Sagnik, et al.
Pubblicazione: (2025)
di: Anupam, Sagnik, et al.
Pubblicazione: (2025)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
di: Yi, Sibo, et al.
Pubblicazione: (2024)
di: Yi, Sibo, et al.
Pubblicazione: (2024)
Adversarial Training Should Be Cast as a Non-Zero-Sum Game
di: Robey, Alexander, et al.
Pubblicazione: (2023)
di: Robey, Alexander, et al.
Pubblicazione: (2023)
Conformal Information Pursuit for Interactively Guiding Large Language Models
di: Chan, Kwan Ho Ryan, et al.
Pubblicazione: (2025)
di: Chan, Kwan Ho Ryan, et al.
Pubblicazione: (2025)
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
di: Cao, Bochuan, et al.
Pubblicazione: (2023)
di: Cao, Bochuan, et al.
Pubblicazione: (2023)
Functional Homotopy: Smoothing Discrete Optimization via Continuous Parameters for LLM Jailbreak Attacks
di: Wang, Zi, et al.
Pubblicazione: (2024)
di: Wang, Zi, et al.
Pubblicazione: (2024)
Watermark Smoothing Attacks against Language Models
di: Chang, Hongyan, et al.
Pubblicazione: (2024)
di: Chang, Hongyan, et al.
Pubblicazione: (2024)
Contextual Safety Reasoning and Grounding for Open-World Robots
di: Ravichandran, Zachary, et al.
Pubblicazione: (2026)
di: Ravichandran, Zachary, et al.
Pubblicazione: (2026)
HSF: Defending against Jailbreak Attacks with Hidden State Filtering
di: Qian, Cheng, et al.
Pubblicazione: (2024)
di: Qian, Cheng, et al.
Pubblicazione: (2024)
Certifying Language Model Robustness with Fuzzed Randomized Smoothing: An Efficient Defense Against Backdoor Attacks
di: He, Bowei, et al.
Pubblicazione: (2025)
di: He, Bowei, et al.
Pubblicazione: (2025)
Be Kind, Rewrite: Benign Projections via Rewriting Defend Against LLM Data Poisoning Attacks
di: Halloran, John T., et al.
Pubblicazione: (2026)
di: Halloran, John T., et al.
Pubblicazione: (2026)
Chordal Sparsity for Lipschitz Constant Estimation of Deep Neural Networks
di: Xue, Anton, et al.
Pubblicazione: (2022)
di: Xue, Anton, et al.
Pubblicazione: (2022)
Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing
di: Zhao, Wei, et al.
Pubblicazione: (2024)
di: Zhao, Wei, et al.
Pubblicazione: (2024)
Defending Against Poisoning Attacks in Federated Learning with Blockchain
di: Dong, Nanqing, et al.
Pubblicazione: (2023)
di: Dong, Nanqing, et al.
Pubblicazione: (2023)
Automated Black-box Prompt Engineering for Personalized Text-to-Image Generation
di: He, Yutong, et al.
Pubblicazione: (2024)
di: He, Yutong, et al.
Pubblicazione: (2024)
Jailbreaking in the Haystack
di: Shah, Rishi Rajesh, et al.
Pubblicazione: (2025)
di: Shah, Rishi Rajesh, et al.
Pubblicazione: (2025)
UniGuard: Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models
di: Oh, Sejoon, et al.
Pubblicazione: (2024)
di: Oh, Sejoon, et al.
Pubblicazione: (2024)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
di: Chu, Junjie, et al.
Pubblicazione: (2024)
di: Chu, Junjie, et al.
Pubblicazione: (2024)
MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation
di: Jiang, Weisen, et al.
Pubblicazione: (2025)
di: Jiang, Weisen, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Jailbreaking Black Box Large Language Models in Twenty Queries
di: Chao, Patrick, et al.
Pubblicazione: (2023) -
Defending Large Language Models against Jailbreak Attacks via Semantic Smoothing
di: Ji, Jiabao, et al.
Pubblicazione: (2024) -
Jailbreaking LLM-Controlled Robots
di: Robey, Alexander, et al.
Pubblicazione: (2024) -
Towards Realistic Guarantees: A Probabilistic Certificate for SmoothLLM
di: Kumarappan, Adarsh, et al.
Pubblicazione: (2025) -
Adversarial Reasoning at Jailbreaking Time
di: Sabbaghi, Mahdi, et al.
Pubblicazione: (2025)