Jailbreaking Black Box Large Language Models in Twenty Queries
Fuente:
arXiv
Saved in:
| Main Authors: | Chao, Patrick, Robey, Alexander, Dobriban, Edgar, Hassani, Hamed, Pappas, George J., Wong, Eric |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
by: Robey, Alexander, et al.
Published: (2023)
by: Robey, Alexander, et al.
Published: (2023)
Evaluating the Performance of Large Language Models via Debates
by: Moniri, Behrad, et al.
Published: (2024)
by: Moniri, Behrad, et al.
Published: (2024)
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
by: Chao, Patrick, et al.
Published: (2024)
by: Chao, Patrick, et al.
Published: (2024)
Conformal Inference under High-Dimensional Covariate Shifts via Likelihood-Ratio Regularization
by: Joshi, Sunay, et al.
Published: (2025)
by: Joshi, Sunay, et al.
Published: (2025)
Conformal Information Pursuit for Interactively Guiding Large Language Models
by: Chan, Kwan Ho Ryan, et al.
Published: (2025)
by: Chan, Kwan Ho Ryan, et al.
Published: (2025)
Adversarial Reasoning at Jailbreaking Time
by: Sabbaghi, Mahdi, et al.
Published: (2025)
by: Sabbaghi, Mahdi, et al.
Published: (2025)
Provable tradeoffs in adversarially robust classification
by: Dobriban, Edgar, et al.
Published: (2020)
by: Dobriban, Edgar, et al.
Published: (2020)
Jailbreaking LLM-Controlled Robots
by: Robey, Alexander, et al.
Published: (2024)
by: Robey, Alexander, et al.
Published: (2024)
Conformal Prediction with Learned Features
by: Kiyani, Shayan, et al.
Published: (2024)
by: Kiyani, Shayan, et al.
Published: (2024)
One-Shot Safety Alignment for Large Language Models via Optimal Dualization
by: Huang, Xinmeng, et al.
Published: (2024)
by: Huang, Xinmeng, et al.
Published: (2024)
Length Optimization in Conformal Prediction
by: Kiyani, Shayan, et al.
Published: (2024)
by: Kiyani, Shayan, et al.
Published: (2024)
Conformal Prediction Beyond the Seen: A Missing Mass Perspective for Uncertainty Quantification in Generative Models
by: Noorani, Sima, et al.
Published: (2025)
by: Noorani, Sima, et al.
Published: (2025)
Decision Theoretic Foundations for Conformal Prediction: Optimal Uncertainty Quantification for Risk-Averse Agents
by: Kiyani, Shayan, et al.
Published: (2025)
by: Kiyani, Shayan, et al.
Published: (2025)
When to Trust the Cheap Check: Weak and Strong Verification for Reasoning
by: Kiyani, Shayan, et al.
Published: (2026)
by: Kiyani, Shayan, et al.
Published: (2026)
Robust Decision Making with Partially Calibrated Forecasts
by: Kiyani, Shayan, et al.
Published: (2025)
by: Kiyani, Shayan, et al.
Published: (2025)
Watermarking Language Models with Error Correcting Codes
by: Chao, Patrick, et al.
Published: (2024)
by: Chao, Patrick, et al.
Published: (2024)
Uncertainty in Language Models: Assessment through Rank-Calibration
by: Huang, Xinmeng, et al.
Published: (2024)
by: Huang, Xinmeng, et al.
Published: (2024)
Temporal Difference Learning with Compressed Updates: Error-Feedback meets Reinforcement Learning
by: Mitra, Aritra, et al.
Published: (2023)
by: Mitra, Aritra, et al.
Published: (2023)
Statistical Methods in Generative AI
by: Dobriban, Edgar
Published: (2025)
by: Dobriban, Edgar
Published: (2025)
Human-AI Collaborative Uncertainty Quantification
by: Noorani, Sima, et al.
Published: (2025)
by: Noorani, Sima, et al.
Published: (2025)
Solving a Research Problem in Mathematical Statistics with AI Assistance
by: Dobriban, Edgar
Published: (2025)
by: Dobriban, Edgar
Published: (2025)
Algorithms for Adversarially Robust Deep Learning
by: Robey, Alexander
Published: (2025)
by: Robey, Alexander
Published: (2025)
Automated Black-box Prompt Engineering for Personalized Text-to-Image Generation
by: He, Yutong, et al.
Published: (2024)
by: He, Yutong, et al.
Published: (2024)
Safety Guardrails for LLM-Enabled Robots
by: Ravichandran, Zachary, et al.
Published: (2025)
by: Ravichandran, Zachary, et al.
Published: (2025)
Benchmarking Misuse Mitigation Against Covert Adversaries
by: Brown, Davis, et al.
Published: (2025)
by: Brown, Davis, et al.
Published: (2025)
Adversarial Training Should Be Cast as a Non-Zero-Sum Game
by: Robey, Alexander, et al.
Published: (2023)
by: Robey, Alexander, et al.
Published: (2023)
BBox-Adapter: Lightweight Adapting for Black-Box Large Language Models
by: Sun, Haotian, et al.
Published: (2024)
by: Sun, Haotian, et al.
Published: (2024)
BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks
by: Anupam, Sagnik, et al.
Published: (2025)
by: Anupam, Sagnik, et al.
Published: (2025)
Jailbreaking in the Haystack
by: Shah, Rishi Rajesh, et al.
Published: (2025)
by: Shah, Rishi Rajesh, et al.
Published: (2025)
Contextual Safety Reasoning and Grounding for Open-World Robots
by: Ravichandran, Zachary, et al.
Published: (2026)
by: Ravichandran, Zachary, et al.
Published: (2026)
Chordal Sparsity for Lipschitz Constant Estimation of Deep Neural Networks
by: Xue, Anton, et al.
Published: (2022)
by: Xue, Anton, et al.
Published: (2022)
BlackDAN: A Black-Box Multi-Objective Approach for Effective and Contextual Jailbreaking of Large Language Models
by: Wang, Xinyuan, et al.
Published: (2024)
by: Wang, Xinyuan, et al.
Published: (2024)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
by: Mehrotra, Anay, et al.
Published: (2023)
by: Mehrotra, Anay, et al.
Published: (2023)
Defending Large Language Models against Jailbreak Attacks via Semantic Smoothing
by: Ji, Jiabao, et al.
Published: (2024)
by: Ji, Jiabao, et al.
Published: (2024)
A Theory of Non-Linear Feature Learning with One Gradient Step in Two-Layer Neural Networks
by: Moniri, Behrad, et al.
Published: (2023)
by: Moniri, Behrad, et al.
Published: (2023)
MultiRisk: Multiple Risk Control via Iterative Score Thresholding
by: Joshi, Sunay, et al.
Published: (2025)
by: Joshi, Sunay, et al.
Published: (2025)
Adversarial Attacks on Robotic Vision Language Action Models
by: Jones, Eliot Krzysztof, et al.
Published: (2025)
by: Jones, Eliot Krzysztof, et al.
Published: (2025)
Adaptively profiling models with task elicitation
by: Brown, Davis, et al.
Published: (2025)
by: Brown, Davis, et al.
Published: (2025)
Large Language Model Confidence Estimation via Black-Box Access
by: Pedapati, Tejaswini, et al.
Published: (2024)
by: Pedapati, Tejaswini, et al.
Published: (2024)
Knowledgeable Language Models as Black-Box Optimizers for Personalized Medicine
by: Yao, Michael S., et al.
Published: (2025)
by: Yao, Michael S., et al.
Published: (2025)
Similar Items
-
SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
by: Robey, Alexander, et al.
Published: (2023) -
Evaluating the Performance of Large Language Models via Debates
by: Moniri, Behrad, et al.
Published: (2024) -
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
by: Chao, Patrick, et al.
Published: (2024) -
Conformal Inference under High-Dimensional Covariate Shifts via Likelihood-Ratio Regularization
by: Joshi, Sunay, et al.
Published: (2025) -
Conformal Information Pursuit for Interactively Guiding Large Language Models
by: Chan, Kwan Ho Ryan, et al.
Published: (2025)