Safety Alignment Should Be Made More Than Just a Few Tokens Deep
Fuente:
arXiv
Saved in:
| Main Authors: | Qi, Xiangyu, Panda, Ashwinee, Lyu, Kaifeng, Ma, Xiao, Roy, Subhrajit, Beirami, Ahmad, Mittal, Prateek, Henderson, Peter |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Safety Alignment Should Be Made More Than Just A Few Attention Heads
by: Huang, Chao, et al.
Published: (2025)
by: Huang, Chao, et al.
Published: (2025)
A New Linear Scaling Rule for Private Adaptive Hyperparameter Optimization
by: Panda, Ashwinee, et al.
Published: (2022)
by: Panda, Ashwinee, et al.
Published: (2022)
Private Fine-tuning of Large Language Models with Zeroth-order Optimization
by: Tang, Xinyu, et al.
Published: (2024)
by: Tang, Xinyu, et al.
Published: (2024)
Privacy Auditing of Large Language Models
by: Panda, Ashwinee, et al.
Published: (2025)
by: Panda, Ashwinee, et al.
Published: (2025)
Teach LLMs to Phish: Stealing Private Information from Language Models
by: Panda, Ashwinee, et al.
Published: (2024)
by: Panda, Ashwinee, et al.
Published: (2024)
AI Risk Management Should Incorporate Both Safety and Security
by: Qi, Xiangyu, et al.
Published: (2024)
by: Qi, Xiangyu, et al.
Published: (2024)
On Evaluating the Durability of Safeguards for Open-Weight LLMs
by: Qi, Xiangyu, et al.
Published: (2024)
by: Qi, Xiangyu, et al.
Published: (2024)
SoK: Honeypots & LLMs, More Than the Sum of Their Parts?
by: Bridges, Robert A., et al.
Published: (2025)
by: Bridges, Robert A., et al.
Published: (2025)
Shadow-Free Membership Inference Attacks: Recommender Systems Are More Vulnerable Than You Thought
by: Chi, Xiaoxiao, et al.
Published: (2024)
by: Chi, Xiaoxiao, et al.
Published: (2024)
Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
Involuntary In-Context Learning: Exploiting Few-Shot Pattern Completion to Bypass Safety Alignment in GPT-5.4
by: Polyakov, Alex, et al.
Published: (2026)
by: Polyakov, Alex, et al.
Published: (2026)
Unsupervised Log Anomaly Detection with Few Unique Tokens
by: Sulc, Antonin, et al.
Published: (2023)
by: Sulc, Antonin, et al.
Published: (2023)
Position: Towards Resilience Against Adversarial Examples
by: Dai, Sihui, et al.
Published: (2024)
by: Dai, Sihui, et al.
Published: (2024)
Defending Against Prompt Injection With a Few DefensiveTokens
by: Chen, Sizhe, et al.
Published: (2025)
by: Chen, Sizhe, et al.
Published: (2025)
Context manipulation attacks : Web agents are susceptible to corrupted memory
by: Patlan, Atharv Singh, et al.
Published: (2025)
by: Patlan, Atharv Singh, et al.
Published: (2025)
Hiding Your Awful Online Choices Made More Efficient and Secure: A New Privacy-Aware Recommender System
by: Mukherjee, Shibam, et al.
Published: (2024)
by: Mukherjee, Shibam, et al.
Published: (2024)
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
by: Wang, Jiongxiao, et al.
Published: (2024)
by: Wang, Jiongxiao, et al.
Published: (2024)
Opening A Pandora's Box: Things You Should Know in the Era of Custom GPTs
by: Tao, Guanhong, et al.
Published: (2023)
by: Tao, Guanhong, et al.
Published: (2023)
"Explain, Don't Just Warn!" -- A Real-Time Framework for Generating Phishing Warnings with Contextual Cues
by: Roy, Sayak Saha, et al.
Published: (2025)
by: Roy, Sayak Saha, et al.
Published: (2025)
Differential Privacy Made Easy
by: Aitsam, Muhammad
Published: (2022)
by: Aitsam, Muhammad
Published: (2022)
When Safety Becomes a Vulnerability: Exploiting LLM Alignment Homogeneity for Transferable Blocking in RAG
by: Li, Junchen, et al.
Published: (2026)
by: Li, Junchen, et al.
Published: (2026)
Matching Ranks Over Probability Yields Truly Deep Safety Alignment
by: Vega, Jason, et al.
Published: (2025)
by: Vega, Jason, et al.
Published: (2025)
PatchDEMUX: A Certifiably Robust Framework for Multi-label Classifiers Against Adversarial Patches
by: Jacob, Dennis, et al.
Published: (2025)
by: Jacob, Dennis, et al.
Published: (2025)
LLM Agents Should Employ Security Principles
by: Zhang, Kaiyuan, et al.
Published: (2025)
by: Zhang, Kaiyuan, et al.
Published: (2025)
Safety Context Injection: Inference-Time Safety Alignment via Static Filtering and Agentic Analysis
by: Xu, Zhenhao, et al.
Published: (2026)
by: Xu, Zhenhao, et al.
Published: (2026)
ReliabilityRAG: Effective and Provably Robust Defense for RAG-based Web-Search
by: Shen, Zeyu, et al.
Published: (2025)
by: Shen, Zeyu, et al.
Published: (2025)
An Iconic Heavy Hitter Algorithm Made Private
by: Holland, Rayne
Published: (2025)
by: Holland, Rayne
Published: (2025)
Byzantine Failures Harm the Generalization of Robust Distributed Learning Algorithms More Than Data Poisoning
by: Boudou, Thomas, et al.
Published: (2025)
by: Boudou, Thomas, et al.
Published: (2025)
Reimagining Safety Alignment with An Image
by: Xia, Yifan, et al.
Published: (2025)
by: Xia, Yifan, et al.
Published: (2025)
Structured Visual Narratives Undermine Safety Alignment in Multimodal Large Language Models
by: Tan, Rui Yang, et al.
Published: (2026)
by: Tan, Rui Yang, et al.
Published: (2026)
Risks to Zero Trust in a Federated Mission Partner Environment
by: Strandell, Keith, et al.
Published: (2022)
by: Strandell, Keith, et al.
Published: (2022)
CLUE-MARK: Watermarking Diffusion Models using CLWE
by: Shehata, Kareem, et al.
Published: (2024)
by: Shehata, Kareem, et al.
Published: (2024)
Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignment
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs
by: Roh, Jaechul, et al.
Published: (2026)
by: Roh, Jaechul, et al.
Published: (2026)
Pinning Is Futile: You Need More Than Local Dependency Versioning to Defend against Supply Chain Attacks
by: He, Hao, et al.
Published: (2025)
by: He, Hao, et al.
Published: (2025)
More to Extract: Discovering MEV by Token Contract Analysis
by: Chen, Jiaqi, et al.
Published: (2026)
by: Chen, Jiaqi, et al.
Published: (2026)
Adversary-Aware DPO: Enhancing Safety Alignment in Vision Language Models via Adversarial Training
by: Weng, Fenghua, et al.
Published: (2025)
by: Weng, Fenghua, et al.
Published: (2025)
Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism via Probabilistically Ablating Refusal Direction
by: Xie, Yuanbo, et al.
Published: (2025)
by: Xie, Yuanbo, et al.
Published: (2025)
You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation
by: Arazzi, Marco, et al.
Published: (2026)
by: Arazzi, Marco, et al.
Published: (2026)
Fuzzerfly Effect: Hardware Fuzzing for Memory Safety
by: Rostami, Mohamadreza, et al.
Published: (2024)
by: Rostami, Mohamadreza, et al.
Published: (2024)
Similar Items
-
Safety Alignment Should Be Made More Than Just A Few Attention Heads
by: Huang, Chao, et al.
Published: (2025) -
A New Linear Scaling Rule for Private Adaptive Hyperparameter Optimization
by: Panda, Ashwinee, et al.
Published: (2022) -
Private Fine-tuning of Large Language Models with Zeroth-order Optimization
by: Tang, Xinyu, et al.
Published: (2024) -
Privacy Auditing of Large Language Models
by: Panda, Ashwinee, et al.
Published: (2025) -
Teach LLMs to Phish: Stealing Private Information from Language Models
by: Panda, Ashwinee, et al.
Published: (2024)