`For Argument's Sake, Show Me How to Harm Myself!': Jailbreaking LLMs in Suicide and Self-Harm Contexts
Fuente:
arXiv
Saved in:
| Main Authors: | Schoene, Annika M, Canca, Cansu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The First Multilingual Model For The Detection of Suicide Texts
by: Zevallos, Rodolfo, et al.
Published: (2024)
by: Zevallos, Rodolfo, et al.
Published: (2024)
Harmful Prompt Laundering: Jailbreaking LLMs with Abductive Styles and Symbolic Encoding
by: Joo, Seongho, et al.
Published: (2025)
by: Joo, Seongho, et al.
Published: (2025)
Harmful Suicide Content Detection
by: Park, Kyumin, et al.
Published: (2024)
by: Park, Kyumin, et al.
Published: (2024)
Lexicography Saves Lives (LSL): Automatically Translating Suicide-Related Language
by: Schoene, Annika Marie, et al.
Published: (2024)
by: Schoene, Annika Marie, et al.
Published: (2024)
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
by: Chan, Yik Siu, et al.
Published: (2025)
by: Chan, Yik Siu, et al.
Published: (2025)
Self-HarmLLM: Can Large Language Model Harm Itself?
by: Kim, Heehwan, et al.
Published: (2025)
by: Kim, Heehwan, et al.
Published: (2025)
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
by: Laine, Rudolf, et al.
Published: (2024)
by: Laine, Rudolf, et al.
Published: (2024)
Survival at Any Cost? LLMs and the Choice Between Self-Preservation and Human Harm
by: Mohamadi, Alireza, et al.
Published: (2025)
by: Mohamadi, Alireza, et al.
Published: (2025)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
by: Yang, Langqi, et al.
Published: (2025)
by: Yang, Langqi, et al.
Published: (2025)
Harm or Humor: A Multimodal, Multilingual Benchmark for Overt and Covert Harmful Humor
by: Sharshar, Ahmed, et al.
Published: (2026)
by: Sharshar, Ahmed, et al.
Published: (2026)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
by: Li, Jing-Jing, et al.
Published: (2026)
by: Li, Jing-Jing, et al.
Published: (2026)
Are LLMs Enough for Hyperpartisan, Fake, Polarized and Harmful Content Detection? Evaluating In-Context Learning vs. Fine-Tuning
by: Maggini, Michele Joshua, et al.
Published: (2025)
by: Maggini, Michele Joshua, et al.
Published: (2025)
HarmTransform: Transforming Explicit Harmful Queries into Stealthy via Multi-Agent Debate
by: Zhu, Shenzhe
Published: (2025)
by: Zhu, Shenzhe
Published: (2025)
Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs
by: Xing, Wenpeng, et al.
Published: (2025)
by: Xing, Wenpeng, et al.
Published: (2025)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation
by: Huang, Tiansheng, et al.
Published: (2024)
by: Huang, Tiansheng, et al.
Published: (2024)
Different Paths to Harmful Compliance: Behavioral Side Effects and Mechanistic Divergence Across LLM Jailbreaks
by: Kabir, Md Rysul, et al.
Published: (2026)
by: Kabir, Md Rysul, et al.
Published: (2026)
Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content
by: Stepanov, Ihor, et al.
Published: (2026)
by: Stepanov, Ihor, et al.
Published: (2026)
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
by: Pandey, Punya Syon, et al.
Published: (2025)
by: Pandey, Punya Syon, et al.
Published: (2025)
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
by: Choi, Sooyung, et al.
Published: (2025)
by: Choi, Sooyung, et al.
Published: (2025)
Towards Comprehensive Detection of Chinese Harmful Memes
by: Lu, Junyu, et al.
Published: (2024)
by: Lu, Junyu, et al.
Published: (2024)
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
by: Sheshadri, Abhay, et al.
Published: (2024)
by: Sheshadri, Abhay, et al.
Published: (2024)
MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language Models
by: Chen, Zixin, et al.
Published: (2025)
by: Chen, Zixin, et al.
Published: (2025)
FairBelief -- Assessing Harmful Beliefs in Language Models
by: Setzu, Mattia, et al.
Published: (2024)
by: Setzu, Mattia, et al.
Published: (2024)
Human-Guided Harm Recovery for Computer Use Agents
by: Li, Christy, et al.
Published: (2026)
by: Li, Christy, et al.
Published: (2026)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
by: Chaudhary, Maheep, et al.
Published: (2025)
by: Chaudhary, Maheep, et al.
Published: (2025)
Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful Beliefs
by: Cheng, Myra, et al.
Published: (2026)
by: Cheng, Myra, et al.
Published: (2026)
StealthGraph: Exposing Domain-Specific Risks in LLMs through Knowledge-Graph-Guided Harmful Prompt Generation
by: Zheng, Huawei, et al.
Published: (2026)
by: Zheng, Huawei, et al.
Published: (2026)
More of the Same: Persistent Representational Harms Under Increased Representation
by: Mickel, Jennifer, et al.
Published: (2025)
by: Mickel, Jennifer, et al.
Published: (2025)
HALF: Harm-Aware LLM Fairness Evaluation Aligned with Deployment
by: Mekky, Ali, et al.
Published: (2025)
by: Mekky, Ali, et al.
Published: (2025)
Beyond Accuracy: An Explainability-Driven Analysis of Harmful Content Detection
by: Dhara, Trishita, et al.
Published: (2026)
by: Dhara, Trishita, et al.
Published: (2026)
Mitigating Harmful Erraticism in LLMs Through Dialectical Behavior Therapy Based De-Escalation Strategies
by: Rangarajan, Pooja, et al.
Published: (2025)
by: Rangarajan, Pooja, et al.
Published: (2025)
Show Me How It's Done: The Role of Explanations in Fine-Tuning Language Models
by: Ballout, Mohamad, et al.
Published: (2024)
by: Ballout, Mohamad, et al.
Published: (2024)
REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM
by: Jindal, Madhur, et al.
Published: (2025)
by: Jindal, Madhur, et al.
Published: (2025)
SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early Alignment
by: Jeung, Wonje, et al.
Published: (2025)
by: Jeung, Wonje, et al.
Published: (2025)
ToxiCraft: A Novel Framework for Synthetic Generation of Harmful Information
by: Hui, Zheng, et al.
Published: (2024)
by: Hui, Zheng, et al.
Published: (2024)
Easier to Mislead Than to Correct: Harmful and Beneficial Revision in LLM Conformity
by: Qu, Jiaming, et al.
Published: (2026)
by: Qu, Jiaming, et al.
Published: (2026)
ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark
by: Liu, Kangwei, et al.
Published: (2025)
by: Liu, Kangwei, et al.
Published: (2025)
MIND: A Multi-agent Framework for Zero-shot Harmful Meme Detection
by: Liu, Ziyan, et al.
Published: (2025)
by: Liu, Ziyan, et al.
Published: (2025)
Lost in Translation? A Comparative Study on the Cross-Lingual Transfer of Composite Harms
by: Shukla, Vaibhav, et al.
Published: (2026)
by: Shukla, Vaibhav, et al.
Published: (2026)
Similar Items
-
The First Multilingual Model For The Detection of Suicide Texts
by: Zevallos, Rodolfo, et al.
Published: (2024) -
Harmful Prompt Laundering: Jailbreaking LLMs with Abductive Styles and Symbolic Encoding
by: Joo, Seongho, et al.
Published: (2025) -
Harmful Suicide Content Detection
by: Park, Kyumin, et al.
Published: (2024) -
Lexicography Saves Lives (LSL): Automatically Translating Suicide-Related Language
by: Schoene, Annika Marie, et al.
Published: (2024) -
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
by: Chan, Yik Siu, et al.
Published: (2025)