Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chan, Yik Siu, Ri, Narutatsu, Xiao, Yuxin, Ghassemi, Marzyeh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment
von: Xiao, Yuxin, et al.
Veröffentlicht: (2025)
von: Xiao, Yuxin, et al.
Veröffentlicht: (2025)
MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-Making
von: Kim, Yubin, et al.
Veröffentlicht: (2024)
von: Kim, Yubin, et al.
Veröffentlicht: (2024)
SFTMix: Elevating Language Model Instruction Tuning with Mixup Recipe
von: Xiao, Yuxin, et al.
Veröffentlicht: (2024)
von: Xiao, Yuxin, et al.
Veröffentlicht: (2024)
In the Name of Fairness: Assessing the Bias in Clinical Record De-identification
von: Xiao, Yuxin, et al.
Veröffentlicht: (2023)
von: Xiao, Yuxin, et al.
Veröffentlicht: (2023)
KScope: A Framework for Characterizing the Knowledge Status of Language Models
von: Xiao, Yuxin, et al.
Veröffentlicht: (2025)
von: Xiao, Yuxin, et al.
Veröffentlicht: (2025)
Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs
von: Naderi, Nariman, et al.
Veröffentlicht: (2025)
von: Naderi, Nariman, et al.
Veröffentlicht: (2025)
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
von: Choi, Sooyung, et al.
Veröffentlicht: (2025)
von: Choi, Sooyung, et al.
Veröffentlicht: (2025)
Can We Predict Alignment Before Models Finish Thinking? Towards Monitoring Misaligned Reasoning Models
von: Chan, Yik Siu, et al.
Veröffentlicht: (2025)
von: Chan, Yik Siu, et al.
Veröffentlicht: (2025)
Laissez-Faire Harms: Algorithmic Biases in Generative Language Models
von: Shieh, Evan, et al.
Veröffentlicht: (2024)
von: Shieh, Evan, et al.
Veröffentlicht: (2024)
Beyond Behaviorist Representational Harms: A Plan for Measurement and Mitigation
von: Chien, Jennifer, et al.
Veröffentlicht: (2024)
von: Chien, Jennifer, et al.
Veröffentlicht: (2024)
Fairer Preferences Elicit Improved Human-Aligned Large Language Model Judgments
von: Zhou, Han, et al.
Veröffentlicht: (2024)
von: Zhou, Han, et al.
Veröffentlicht: (2024)
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
von: Gringras, David
Veröffentlicht: (2026)
von: Gringras, David
Veröffentlicht: (2026)
Reaching Beyond the Mode: RL for Distributional Reasoning in Language Models
von: Puri, Isha, et al.
Veröffentlicht: (2026)
von: Puri, Isha, et al.
Veröffentlicht: (2026)
The Impact of Large Language Models in Academia: from Writing to Speaking
von: Geng, Mingmeng, et al.
Veröffentlicht: (2024)
von: Geng, Mingmeng, et al.
Veröffentlicht: (2024)
When Correct Beliefs Collapse: Epistemic Resilience of LLMs under Clinical Pressure
von: Xiao, Boyu, et al.
Veröffentlicht: (2026)
von: Xiao, Boyu, et al.
Veröffentlicht: (2026)
Generative AI in Medicine
von: Shanmugam, Divya, et al.
Veröffentlicht: (2024)
von: Shanmugam, Divya, et al.
Veröffentlicht: (2024)
Identifying Implicit Social Biases in Vision-Language Models
von: Hamidieh, Kimia, et al.
Veröffentlicht: (2024)
von: Hamidieh, Kimia, et al.
Veröffentlicht: (2024)
Train Long, Think Short: Curriculum Learning for Efficient Reasoning
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2025)
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2025)
Easy Problems That LLMs Get Wrong
von: Williams, Sean, et al.
Veröffentlicht: (2024)
von: Williams, Sean, et al.
Veröffentlicht: (2024)
Wikipedia in the Era of LLMs: Evolution and Risks
von: Huang, Siming, et al.
Veröffentlicht: (2025)
von: Huang, Siming, et al.
Veröffentlicht: (2025)
"They are uncultured": Unveiling Covert Harms and Social Threats in LLM Generated Conversations
von: Dammu, Preetam Prabhu Srikar, et al.
Veröffentlicht: (2024)
von: Dammu, Preetam Prabhu Srikar, et al.
Veröffentlicht: (2024)
Reasoning-Based Refinement of Unsupervised Text Clusters with LLMs
von: Islam, Tunazzina
Veröffentlicht: (2026)
von: Islam, Tunazzina
Veröffentlicht: (2026)
Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations
von: Hamman, Faisal, et al.
Veröffentlicht: (2025)
von: Hamman, Faisal, et al.
Veröffentlicht: (2025)
Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases
von: AlDahoul, Nouar, et al.
Veröffentlicht: (2025)
von: AlDahoul, Nouar, et al.
Veröffentlicht: (2025)
Towards Robust Evaluation of Unlearning in LLMs via Data Transformations
von: Joshi, Abhinav, et al.
Veröffentlicht: (2024)
von: Joshi, Abhinav, et al.
Veröffentlicht: (2024)
How Far Are We From AGI: Are LLMs All We Need?
von: Feng, Tao, et al.
Veröffentlicht: (2024)
von: Feng, Tao, et al.
Veröffentlicht: (2024)
The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs
von: Han, Pengrui, et al.
Veröffentlicht: (2025)
von: Han, Pengrui, et al.
Veröffentlicht: (2025)
Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content
von: Stepanov, Ihor, et al.
Veröffentlicht: (2026)
von: Stepanov, Ihor, et al.
Veröffentlicht: (2026)
Evaluating Reasoning LLMs for Suicide Screening with the Columbia-Suicide Severity Rating Scale
von: Patil, Avinash, et al.
Veröffentlicht: (2025)
von: Patil, Avinash, et al.
Veröffentlicht: (2025)
We're Different, We're the Same: Creative Homogeneity Across LLMs
von: Wenger, Emily, et al.
Veröffentlicht: (2025)
von: Wenger, Emily, et al.
Veröffentlicht: (2025)
From Perceptions to Decisions: Wildfire Evacuation Decision Prediction with Behavioral Theory-informed LLMs
von: Chen, Ruxiao, et al.
Veröffentlicht: (2025)
von: Chen, Ruxiao, et al.
Veröffentlicht: (2025)
Discourse vs emissions: Analysis of corporate narratives, symbolic practices, and mimicry through LLMs
von: Hassani, Bertrand Kian, et al.
Veröffentlicht: (2025)
von: Hassani, Bertrand Kian, et al.
Veröffentlicht: (2025)
Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions
von: Huang, Saffron, et al.
Veröffentlicht: (2025)
von: Huang, Saffron, et al.
Veröffentlicht: (2025)
ArzEn-LLM: Code-Switched Egyptian Arabic-English Translation and Speech Recognition Using LLMs
von: Heakl, Ahmed, et al.
Veröffentlicht: (2024)
von: Heakl, Ahmed, et al.
Veröffentlicht: (2024)
Sustainable LLM Inference for Edge AI: Evaluating Quantized LLMs for Energy Efficiency, Output Accuracy, and Inference Latency
von: Husom, Erik Johannes, et al.
Veröffentlicht: (2025)
von: Husom, Erik Johannes, et al.
Veröffentlicht: (2025)
MateInfoUB: A Real-World Benchmark for Testing LLMs in Competitive, Multilingual, and Multimodal Educational Tasks
von: Marius, Dumitran Adrian, et al.
Veröffentlicht: (2025)
von: Marius, Dumitran Adrian, et al.
Veröffentlicht: (2025)
Efficient Safety Retrofitting Against Jailbreaking for LLMs
von: Garcia-Gasulla, Dario, et al.
Veröffentlicht: (2025)
von: Garcia-Gasulla, Dario, et al.
Veröffentlicht: (2025)
COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability
von: Guo, Xingang, et al.
Veröffentlicht: (2024)
von: Guo, Xingang, et al.
Veröffentlicht: (2024)
BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Elicitation
von: Islam, Sekh Mainul, et al.
Veröffentlicht: (2025)
von: Islam, Sekh Mainul, et al.
Veröffentlicht: (2025)
All in How You Ask for It: Simple Black-Box Method for Jailbreak Attacks
von: Takemoto, Kazuhiro
Veröffentlicht: (2024)
von: Takemoto, Kazuhiro
Veröffentlicht: (2024)
Ähnliche Einträge
-
When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment
von: Xiao, Yuxin, et al.
Veröffentlicht: (2025) -
MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-Making
von: Kim, Yubin, et al.
Veröffentlicht: (2024) -
SFTMix: Elevating Language Model Instruction Tuning with Mixup Recipe
von: Xiao, Yuxin, et al.
Veröffentlicht: (2024) -
In the Name of Fairness: Assessing the Bias in Clinical Record De-identification
von: Xiao, Yuxin, et al.
Veröffentlicht: (2023) -
KScope: A Framework for Characterizing the Knowledge Status of Language Models
von: Xiao, Yuxin, et al.
Veröffentlicht: (2025)