How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhou, Zhenhong, Yu, Haiyang, Zhang, Xinghua, Xu, Rongwu, Huang, Fei, Li, Yongbin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
On the Role of Attention Heads in Large Language Model Safety
di: Zhou, Zhenhong, et al.
Pubblicazione: (2024)
di: Zhou, Zhenhong, et al.
Pubblicazione: (2024)
Nuclear Deployed: Analyzing Catastrophic Risks in Decision-making of Autonomous LLM Agents
di: Xu, Rongwu, et al.
Pubblicazione: (2025)
di: Xu, Rongwu, et al.
Pubblicazione: (2025)
The Shawshank Redemption of Embodied AI: Understanding and Benchmarking Indirect Environmental Jailbreaks
di: Li, Chunyang, et al.
Pubblicazione: (2025)
di: Li, Chunyang, et al.
Pubblicazione: (2025)
How the Future Works at SOUPS: Analyzing Future Work Statements and Their Impact on Usable Security and Privacy Research
di: Suray, Jacques, et al.
Pubblicazione: (2024)
di: Suray, Jacques, et al.
Pubblicazione: (2024)
How Safe Is Your Data in Connected and Autonomous Cars: A Consumer Advantage or a Privacy Nightmare ?
di: Chougule, Amit, et al.
Pubblicazione: (2026)
di: Chougule, Amit, et al.
Pubblicazione: (2026)
LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
di: Li, Nathaniel, et al.
Pubblicazione: (2024)
di: Li, Nathaniel, et al.
Pubblicazione: (2024)
LightDefense: A Lightweight Uncertainty-Driven Defense against Jailbreaks via Shifted Token Distribution
di: Yang, Zhuoran, et al.
Pubblicazione: (2025)
di: Yang, Zhuoran, et al.
Pubblicazione: (2025)
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
di: Murphy, Brendan, et al.
Pubblicazione: (2025)
di: Murphy, Brendan, et al.
Pubblicazione: (2025)
SafeAligner: Safety Alignment against Jailbreak Attacks via Response Disparity Guidance
di: Huang, Caishuang, et al.
Pubblicazione: (2024)
di: Huang, Caishuang, et al.
Pubblicazione: (2024)
Data Traceability for Privacy Alignment
di: Liao, Kevin, et al.
Pubblicazione: (2025)
di: Liao, Kevin, et al.
Pubblicazione: (2025)
Careful About What App Promotion Ads Recommend! Detecting and Explaining Malware Promotion via App Promotion Graph
di: Ma, Shang, et al.
Pubblicazione: (2024)
di: Ma, Shang, et al.
Pubblicazione: (2024)
An Evaluation of Chat Safety Moderations in Roblox
di: Kaushik, Priya, et al.
Pubblicazione: (2026)
di: Kaushik, Priya, et al.
Pubblicazione: (2026)
The Earth is Flat because...: Investigating LLMs' Belief towards Misinformation via Persuasive Conversation
di: Xu, Rongwu, et al.
Pubblicazione: (2023)
di: Xu, Rongwu, et al.
Pubblicazione: (2023)
How Generative AI Empowers Attackers and Defenders Across the Trust & Safety Landscape
di: Kelley, Patrick Gage, et al.
Pubblicazione: (2025)
di: Kelley, Patrick Gage, et al.
Pubblicazione: (2025)
Safety case template for frontier AI: A cyber inability argument
di: Goemans, Arthur, et al.
Pubblicazione: (2024)
di: Goemans, Arthur, et al.
Pubblicazione: (2024)
Superficial Safety Alignment Hypothesis
di: Li, Jianwei, et al.
Pubblicazione: (2024)
di: Li, Jianwei, et al.
Pubblicazione: (2024)
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures
di: Zhou, Yukai, et al.
Pubblicazione: (2025)
di: Zhou, Yukai, et al.
Pubblicazione: (2025)
Proactive defense against LLM Jailbreak
di: Zhao, Weiliang, et al.
Pubblicazione: (2025)
di: Zhao, Weiliang, et al.
Pubblicazione: (2025)
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
di: Wang, Jiongxiao, et al.
Pubblicazione: (2024)
di: Wang, Jiongxiao, et al.
Pubblicazione: (2024)
Advancing Highway Work Zone Safety: A Comprehensive Review of Sensor Technologies for Intrusion and Proximity Hazards
di: Demeke, Ayenew Yihune, et al.
Pubblicazione: (2025)
di: Demeke, Ayenew Yihune, et al.
Pubblicazione: (2025)
The Shifting Landscape of Cybersecurity: The Impact of Remote Work and COVID-19 on Data Breach Trends
di: Ozer, Murat, et al.
Pubblicazione: (2024)
di: Ozer, Murat, et al.
Pubblicazione: (2024)
Evaluating the Critical Risks of Amazon's Nova Premier under the Frontier Model Safety Framework
di: Krishna, Satyapriya, et al.
Pubblicazione: (2025)
di: Krishna, Satyapriya, et al.
Pubblicazione: (2025)
Safety and Security Analysis of Large Language Models: Benchmarking Risk Profile and Harm Potential
di: Akiri, Charankumar, et al.
Pubblicazione: (2025)
di: Akiri, Charankumar, et al.
Pubblicazione: (2025)
Integrating Generative AI into Cybersecurity Education: A Study of OCR and Multimodal LLM-assisted Instruction
di: Patel, Karan, et al.
Pubblicazione: (2025)
di: Patel, Karan, et al.
Pubblicazione: (2025)
Pwned: How Often Are Americans' Online Accounts Breached?
di: Cor, Ken, et al.
Pubblicazione: (2018)
di: Cor, Ken, et al.
Pubblicazione: (2018)
Firewalls to Secure Dynamic LLM Agentic Networks
di: Abdelnabi, Sahar, et al.
Pubblicazione: (2025)
di: Abdelnabi, Sahar, et al.
Pubblicazione: (2025)
How to Manage My Data? With Machine--Interpretable GDPR Rights!
di: Esteves, Beatriz, et al.
Pubblicazione: (2024)
di: Esteves, Beatriz, et al.
Pubblicazione: (2024)
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
di: Shen, Guobin, et al.
Pubblicazione: (2025)
di: Shen, Guobin, et al.
Pubblicazione: (2025)
Data Protection through Governance Frameworks
di: Julakanti, Sivananda Reddy, et al.
Pubblicazione: (2025)
di: Julakanti, Sivananda Reddy, et al.
Pubblicazione: (2025)
Case Studies: Effective Approaches for Navigating Cross-Border Cloud Data Transfers Amid U.S. Government Privacy and Safety Concerns
di: Adebayo, Motunrayo
Pubblicazione: (2025)
di: Adebayo, Motunrayo
Pubblicazione: (2025)
How Jailbreak Defenses Work and Ensemble? A Mechanistic Investigation
di: Long, Zhuohang, et al.
Pubblicazione: (2025)
di: Long, Zhuohang, et al.
Pubblicazione: (2025)
Sark: Oblivious Integrity Without Global State
di: Lynham, Alex, et al.
Pubblicazione: (2025)
di: Lynham, Alex, et al.
Pubblicazione: (2025)
ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations
di: Gomaa, Amr, et al.
Pubblicazione: (2025)
di: Gomaa, Amr, et al.
Pubblicazione: (2025)
You Still See Me: How Data Protection Supports the Architecture of AI Surveillance
di: Yew, Rui-Jie, et al.
Pubblicazione: (2024)
di: Yew, Rui-Jie, et al.
Pubblicazione: (2024)
TombRaider: Entering the Vault of History to Jailbreak Large Language Models
di: Ding, Junchen, et al.
Pubblicazione: (2025)
di: Ding, Junchen, et al.
Pubblicazione: (2025)
Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking
di: Wu, Yu-Hang, et al.
Pubblicazione: (2025)
di: Wu, Yu-Hang, et al.
Pubblicazione: (2025)
When Your Reviewer is an LLM: Biases, Divergence, and Prompt Injection Risks in Peer Review
di: Zhu, Changjia, et al.
Pubblicazione: (2025)
di: Zhu, Changjia, et al.
Pubblicazione: (2025)
Farewell to Westphalia: Crypto Sovereignty and Post-Nation-State Governaance
di: Hope, Jarrad, et al.
Pubblicazione: (2025)
di: Hope, Jarrad, et al.
Pubblicazione: (2025)
Risks & Benefits of LLMs & GenAI for Platform Integrity, Healthcare Diagnostics, Financial Trust and Compliance, Cybersecurity, Privacy & AI Safety: A Comprehensive Survey, Roadmap & Implementation Blueprint
di: Ahi, Kiarash
Pubblicazione: (2025)
di: Ahi, Kiarash
Pubblicazione: (2025)
adF: A Novel System for Measuring Web Fingerprinting through Ads
di: Bermejo-Agueda, Miguel A., et al.
Pubblicazione: (2023)
di: Bermejo-Agueda, Miguel A., et al.
Pubblicazione: (2023)
Documenti analoghi
-
On the Role of Attention Heads in Large Language Model Safety
di: Zhou, Zhenhong, et al.
Pubblicazione: (2024) -
Nuclear Deployed: Analyzing Catastrophic Risks in Decision-making of Autonomous LLM Agents
di: Xu, Rongwu, et al.
Pubblicazione: (2025) -
The Shawshank Redemption of Embodied AI: Understanding and Benchmarking Indirect Environmental Jailbreaks
di: Li, Chunyang, et al.
Pubblicazione: (2025) -
How the Future Works at SOUPS: Analyzing Future Work Statements and Their Impact on Usable Security and Privacy Research
di: Suray, Jacques, et al.
Pubblicazione: (2024) -
How Safe Is Your Data in Connected and Autonomous Cars: A Consumer Advantage or a Privacy Nightmare ?
di: Chougule, Amit, et al.
Pubblicazione: (2026)