An Audit and Analysis of LLM-Assisted Health Misinformation Jailbreaks Against LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Hussain, Ayana, Zhao, Patrick, Vincent, Nicholas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
by: Liu, Fan, et al.
Published: (2024)
by: Liu, Fan, et al.
Published: (2024)
LLM Robustness Against Misinformation in Biomedical Question Answering
by: Bondarenko, Alexander, et al.
Published: (2024)
by: Bondarenko, Alexander, et al.
Published: (2024)
Beyond Jailbreaking: Auditing Contextual Privacy in LLM Agents
by: Das, Saswat, et al.
Published: (2025)
by: Das, Saswat, et al.
Published: (2025)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
by: Chu, Junjie, et al.
Published: (2024)
by: Chu, Junjie, et al.
Published: (2024)
FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks
by: Chen, Bocheng, et al.
Published: (2024)
by: Chen, Bocheng, et al.
Published: (2024)
Efficient Safety Retrofitting Against Jailbreaking for LLMs
by: Garcia-Gasulla, Dario, et al.
Published: (2025)
by: Garcia-Gasulla, Dario, et al.
Published: (2025)
JailNewsBench: Multi-Lingual and Regional Benchmark for Fake News Generation under Jailbreak Attacks
by: Kaneko, Masahiro, et al.
Published: (2026)
by: Kaneko, Masahiro, et al.
Published: (2026)
Multi-Agent Retrieval-Augmented Framework for Evidence-Based Counterspeech Against Health Misinformation
by: Anik, Anirban Saha, et al.
Published: (2025)
by: Anik, Anirban Saha, et al.
Published: (2025)
Semantic Mirror Jailbreak: Genetic Algorithm Based Jailbreak Prompts Against Open-source LLMs
by: Li, Xiaoxia, et al.
Published: (2024)
by: Li, Xiaoxia, et al.
Published: (2024)
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification
by: Li, Yu, et al.
Published: (2025)
by: Li, Yu, et al.
Published: (2025)
JailbreakLens: Visual Analysis of Jailbreak Attacks Against Large Language Models
by: Feng, Yingchaojie, et al.
Published: (2024)
by: Feng, Yingchaojie, et al.
Published: (2024)
Rectifying Belief Space via Unlearning to Harness LLMs' Reasoning
by: Niwa, Ayana, et al.
Published: (2025)
by: Niwa, Ayana, et al.
Published: (2025)
Emoji Attack: Enhancing Jailbreak Attacks Against Judge LLM Detection
by: Wei, Zhipeng, et al.
Published: (2024)
by: Wei, Zhipeng, et al.
Published: (2024)
Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers
by: Lin, Liang, et al.
Published: (2025)
by: Lin, Liang, et al.
Published: (2025)
Epistemic Blinding: An Inference-Time Protocol for Auditing Prior Contamination in LLM-Assisted Analysis
by: Cuccarese, Michael
Published: (2026)
by: Cuccarese, Michael
Published: (2026)
Offscript: Automated Auditing of Instruction Adherence in LLMs
by: Clark, Nicholas, et al.
Published: (2025)
by: Clark, Nicholas, et al.
Published: (2025)
Assisted Counterspeech Writing at the Crossroads of Hate Speech and Misinformation
by: Martone, Genoveffa, et al.
Published: (2026)
by: Martone, Genoveffa, et al.
Published: (2026)
Intention Analysis Makes LLMs A Good Jailbreak Defender
by: Zhang, Yuqi, et al.
Published: (2024)
by: Zhang, Yuqi, et al.
Published: (2024)
Unraveling Misinformation Propagation in LLM Reasoning
by: Feng, Yiyang, et al.
Published: (2025)
by: Feng, Yiyang, et al.
Published: (2025)
Efficient LLM-Jailbreaking via Multimodal-LLM Jailbreak
by: Ji, Haoxuan, et al.
Published: (2024)
by: Ji, Haoxuan, et al.
Published: (2024)
Beyond the Crowd: LLM-Augmented Community Notes for Governing Health Misinformation
by: Wu, Jiaying, et al.
Published: (2025)
by: Wu, Jiaying, et al.
Published: (2025)
Towards Understanding Jailbreak Attacks in LLMs: A Representation Space Analysis
by: Lin, Yuping, et al.
Published: (2024)
by: Lin, Yuping, et al.
Published: (2024)
Misinforming LLMs: vulnerabilities, challenges and opportunities
by: Zhou, Bo, et al.
Published: (2024)
by: Zhou, Bo, et al.
Published: (2024)
When Cow Urine Cures Constipation on YouTube: Limits of LLMs in Detecting Culture-specific Health Misinformation
by: Khan, Anamta, et al.
Published: (2026)
by: Khan, Anamta, et al.
Published: (2026)
Proactive defense against LLM Jailbreak
by: Zhao, Weiliang, et al.
Published: (2025)
by: Zhao, Weiliang, et al.
Published: (2025)
AttnGCG: Enhancing Jailbreaking Attacks on LLMs with Attention Manipulation
by: Wang, Zijun, et al.
Published: (2024)
by: Wang, Zijun, et al.
Published: (2024)
Machine Learning for Detection and Analysis of Novel LLM Jailbreaks
by: Hawkins, John, et al.
Published: (2025)
by: Hawkins, John, et al.
Published: (2025)
Unmasking Digital Falsehoods: A Comparative Analysis of LLM-Based Misinformation Detection Strategies
by: Huang, Tianyi, et al.
Published: (2025)
by: Huang, Tianyi, et al.
Published: (2025)
xJailbreak: Representation Space Guided Reinforcement Learning for Interpretable LLM Jailbreaking
by: Lee, Sunbowen, et al.
Published: (2025)
by: Lee, Sunbowen, et al.
Published: (2025)
Exploring the Impact of Instruction-Tuning on LLM's Susceptibility to Misinformation
by: Han, Kyubeen, et al.
Published: (2025)
by: Han, Kyubeen, et al.
Published: (2025)
Learn and Unlearn: Addressing Misinformation in Multilingual LLMs
by: Lu, Taiming, et al.
Published: (2024)
by: Lu, Taiming, et al.
Published: (2024)
Understanding Knowledge Drift in LLMs through Misinformation
by: Fastowski, Alina, et al.
Published: (2024)
by: Fastowski, Alina, et al.
Published: (2024)
Evaluating the Simulation of Human Personality-Driven Susceptibility to Misinformation with LLMs
by: Pratelli, Manuel, et al.
Published: (2025)
by: Pratelli, Manuel, et al.
Published: (2025)
Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs
by: Xu, Zhao, et al.
Published: (2024)
by: Xu, Zhao, et al.
Published: (2024)
[WIP] Jailbreak Paradox: The Achilles' Heel of LLMs
by: Rao, Abhinav, et al.
Published: (2024)
by: Rao, Abhinav, et al.
Published: (2024)
RAEmoLLM: Retrieval Augmented LLMs for Cross-Domain Misinformation Detection Using In-Context Learning Based on Emotional Information
by: Liu, Zhiwei, et al.
Published: (2024)
by: Liu, Zhiwei, et al.
Published: (2024)
AmbigNLG: Addressing Task Ambiguity in Instruction for NLG
by: Niwa, Ayana, et al.
Published: (2024)
by: Niwa, Ayana, et al.
Published: (2024)
Analyzing the Inherent Response Tendency of LLMs: Real-World Instructions-Driven Jailbreak
by: Du, Yanrui, et al.
Published: (2023)
by: Du, Yanrui, et al.
Published: (2023)
Merging Improves Self-Critique Against Jailbreak Attacks
by: Gallego, Victor
Published: (2024)
by: Gallego, Victor
Published: (2024)
Explore the Potential of LLMs in Misinformation Detection: An Empirical Study
by: Chen, Mengyang, et al.
Published: (2023)
by: Chen, Mengyang, et al.
Published: (2023)
Similar Items
-
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
by: Liu, Fan, et al.
Published: (2024) -
LLM Robustness Against Misinformation in Biomedical Question Answering
by: Bondarenko, Alexander, et al.
Published: (2024) -
Beyond Jailbreaking: Auditing Contextual Privacy in LLM Agents
by: Das, Saswat, et al.
Published: (2025) -
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
by: Chu, Junjie, et al.
Published: (2024) -
FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks
by: Chen, Bocheng, et al.
Published: (2024)