All in How You Ask for It: Simple Black-Box Method for Jailbreak Attacks
Fuente:
arXiv
Saved in:
| Main Author: | Takemoto, Kazuhiro |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
by: Chan, Yik Siu, et al.
Published: (2025)
by: Chan, Yik Siu, et al.
Published: (2025)
Text Corpora as Concept Fields: Black-Box Hallucination and Novelty Measurement
by: Kersting, Nicholas S., et al.
Published: (2026)
by: Kersting, Nicholas S., et al.
Published: (2026)
Graph of Attacks: Improved Black-Box and Interpretable Jailbreaks for LLMs
by: Akbar-Tajari, Mohammad, et al.
Published: (2025)
by: Akbar-Tajari, Mohammad, et al.
Published: (2025)
Benchmarking the Safety of Large Language Models for Robotic Health Attendant Control
by: Nakao, Mahiro, et al.
Published: (2026)
by: Nakao, Mahiro, et al.
Published: (2026)
Effective and Efficient Jailbreaks of Black-Box LLMs with Cross-Behavior Attacks
by: Gohil, Vasudev
Published: (2025)
by: Gohil, Vasudev
Published: (2025)
The Moral Machine Experiment on Large Language Models
by: Takemoto, Kazuhiro
Published: (2023)
by: Takemoto, Kazuhiro
Published: (2023)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
by: Mehrotra, Anay, et al.
Published: (2023)
by: Mehrotra, Anay, et al.
Published: (2023)
On Wednesdays, We Ask Questions: Optimizing "Active Listening" in Automated Legal Triage and Referral
by: Steenhuis, Quinten, et al.
Published: (2026)
by: Steenhuis, Quinten, et al.
Published: (2026)
Scaling Law in LLM Simulated Personality: More Detailed and Realistic Persona Profile Is All You Need
by: Bai, Yuqi, et al.
Published: (2025)
by: Bai, Yuqi, et al.
Published: (2025)
From Black-Box Confidence to Measurable Trust in Clinical AI: A Framework for Evidence, Supervision, and Staged Autonomy
by: Zabolotnii, Serhii, et al.
Published: (2026)
by: Zabolotnii, Serhii, et al.
Published: (2026)
Ask LLMs Directly, "What shapes your bias?": Measuring Social Bias in Large Language Models
by: Shin, Jisu, et al.
Published: (2024)
by: Shin, Jisu, et al.
Published: (2024)
Machine Learning for Detection and Analysis of Novel LLM Jailbreaks
by: Hawkins, John, et al.
Published: (2025)
by: Hawkins, John, et al.
Published: (2025)
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
by: Murphy, Brendan, et al.
Published: (2025)
by: Murphy, Brendan, et al.
Published: (2025)
How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
by: Zhou, Zhenhong, et al.
Published: (2024)
by: Zhou, Zhenhong, et al.
Published: (2024)
Is Contrasting All You Need? Contrastive Learning for the Detection and Attribution of AI-generated Text
by: La Cava, Lucio, et al.
Published: (2024)
by: La Cava, Lucio, et al.
Published: (2024)
Mass-Scale Analysis of In-the-Wild Conversations Reveals Complexity Bounds on LLM Jailbreaking
by: Creo, Aldan, et al.
Published: (2025)
by: Creo, Aldan, et al.
Published: (2025)
Unmasking the Canvas: A Dynamic Benchmark for Image Generation Jailbreaking and LLM Content Safety
by: Nair, Variath Madhupal Gautham, et al.
Published: (2025)
by: Nair, Variath Madhupal Gautham, et al.
Published: (2025)
What Do LLMs Associate with Your Name? A Human-Centered Black-Box Audit of Personal Data
by: Staufer, Dimitri, et al.
Published: (2026)
by: Staufer, Dimitri, et al.
Published: (2026)
How Far Are We From AGI: Are LLMs All We Need?
by: Feng, Tao, et al.
Published: (2024)
by: Feng, Tao, et al.
Published: (2024)
Language Models Change Facts Based on the Way You Talk
by: Kearney, Matthew, et al.
Published: (2025)
by: Kearney, Matthew, et al.
Published: (2025)
"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks
by: Wang, Libo
Published: (2024)
by: Wang, Libo
Published: (2024)
Workflow is All You Need: Escaping the "Statistical Smoothing Trap" via High-Entropy Information Foraging and Adversarial Pacing
by: Jiang, Zhongjie
Published: (2025)
by: Jiang, Zhongjie
Published: (2025)
TombRaider: Entering the Vault of History to Jailbreak Large Language Models
by: Ding, Junchen, et al.
Published: (2025)
by: Ding, Junchen, et al.
Published: (2025)
Compromising Honesty and Harmlessness in Language Models via Deception Attacks
by: Vaugrante, Laurène, et al.
Published: (2025)
by: Vaugrante, Laurène, et al.
Published: (2025)
Truth Sleuth and Trend Bender: AI Agents to fact-check YouTube videos and influence opinions
by: Logé, Cécile, et al.
Published: (2025)
by: Logé, Cécile, et al.
Published: (2025)
Magic, Madness, Heaven, Sin: LLM Output Diversity is Everything, Everywhere, All at Once
by: Dhingra, Harnoor
Published: (2026)
by: Dhingra, Harnoor
Published: (2026)
You've Changed: Detecting Modification of Black-Box Large Language Models
by: Dima, Alden, et al.
Published: (2025)
by: Dima, Alden, et al.
Published: (2025)
Stars, Stripes, and Silicon: Unravelling the ChatGPT's All-American, Monochrome, Cis-centric Bias
by: Torrielli, Federico
Published: (2024)
by: Torrielli, Federico
Published: (2024)
HQA-Attack: Toward High Quality Black-Box Hard-Label Adversarial Attack on Text
by: Liu, Han, et al.
Published: (2024)
by: Liu, Han, et al.
Published: (2024)
How Large Language Models are Designed to Hallucinate
by: Ackermann, Richard, et al.
Published: (2025)
by: Ackermann, Richard, et al.
Published: (2025)
TrailBlazer: History-Guided Reinforcement Learning for Black-Box LLM Jailbreaking
by: Yoon, Sung-Hoon, et al.
Published: (2026)
by: Yoon, Sung-Hoon, et al.
Published: (2026)
What You See is What You Ask: Evaluating Audio Descriptions
by: Kala, Divy, et al.
Published: (2025)
by: Kala, Divy, et al.
Published: (2025)
How Did We Get Here? Summarizing Conversation Dynamics
by: Hua, Yilun, et al.
Published: (2024)
by: Hua, Yilun, et al.
Published: (2024)
The Last Fingerprint: How Markdown Training Shapes LLM Prose
by: Freeburg, E. M.
Published: (2026)
by: Freeburg, E. M.
Published: (2026)
The Incomplete Bridge: How AI Research (Mis)Engages with Psychology
by: Jiang, Han, et al.
Published: (2025)
by: Jiang, Han, et al.
Published: (2025)
Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?
by: Mu, Junjie, et al.
Published: (2025)
by: Mu, Junjie, et al.
Published: (2025)
How English Print Media Frames Human-Elephant Conflicts in India
by: Punith, Bonala Sai, et al.
Published: (2026)
by: Punith, Bonala Sai, et al.
Published: (2026)
White Men Lead, Black Women Help? Benchmarking and Mitigating Language Agency Social Biases in LLMs
by: Wan, Yixin, et al.
Published: (2024)
by: Wan, Yixin, et al.
Published: (2024)
Artificial Intelligence in Brazilian News: A Mixed-Methods Analysis
by: Hernandes, Raphael, et al.
Published: (2024)
by: Hernandes, Raphael, et al.
Published: (2024)
The Silent Curriculum: How Does LLM Monoculture Shape Educational Content and Its Accessibility?
by: Priyanshu, Aman, et al.
Published: (2024)
by: Priyanshu, Aman, et al.
Published: (2024)
Similar Items
-
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
by: Chan, Yik Siu, et al.
Published: (2025) -
Text Corpora as Concept Fields: Black-Box Hallucination and Novelty Measurement
by: Kersting, Nicholas S., et al.
Published: (2026) -
Graph of Attacks: Improved Black-Box and Interpretable Jailbreaks for LLMs
by: Akbar-Tajari, Mohammad, et al.
Published: (2025) -
Benchmarking the Safety of Large Language Models for Robotic Health Attendant Control
by: Nakao, Mahiro, et al.
Published: (2026) -
Effective and Efficient Jailbreaks of Black-Box LLMs with Cross-Behavior Attacks
by: Gohil, Vasudev
Published: (2025)