UK AISI Alignment Evaluation Case-Study
Fuente:
arXiv
Saved in:
| Main Authors: | Souly, Alexandra, Kirk, Robert, Merizian, Jacob, D'Cruz, Abby, Davies, Xander |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating whether AI models would sabotage AI safety research
by: Kirk, Robert, et al.
Published: (2026)
by: Kirk, Robert, et al.
Published: (2026)
Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents
by: Bazinska, Julia, et al.
Published: (2025)
by: Bazinska, Julia, et al.
Published: (2025)
Fundamental Limitations in Pointwise Defences of LLM Finetuning APIs
by: Davies, Xander, et al.
Published: (2025)
by: Davies, Xander, et al.
Published: (2025)
SeCodePLT: A Unified Platform for Evaluating the Security of Code GenAI
by: Nie, Yuzhou, et al.
Published: (2024)
by: Nie, Yuzhou, et al.
Published: (2024)
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
by: Che, Zora, et al.
Published: (2025)
by: Che, Zora, et al.
Published: (2025)
Reimagining Safety Alignment with An Image
by: Xia, Yifan, et al.
Published: (2025)
by: Xia, Yifan, et al.
Published: (2025)
Generative AI-enabled Blockchain Networks: Fundamentals, Applications, and Case Study
by: Nguyen, Cong T., et al.
Published: (2024)
by: Nguyen, Cong T., et al.
Published: (2024)
Artificial Intelligence as the New Hacker: Developing Agents for Offensive Security
by: Valencia, Leroy Jacob
Published: (2024)
by: Valencia, Leroy Jacob
Published: (2024)
EVA: Editing for Versatile Alignment against Jailbreaks
by: Wang, Yi, et al.
Published: (2026)
by: Wang, Yi, et al.
Published: (2026)
Agent Safety Alignment via Reinforcement Learning
by: Sha, Zeyang, et al.
Published: (2025)
by: Sha, Zeyang, et al.
Published: (2025)
Supervised and Unsupervised Alignments for Spoofing Behavioral Biometrics
by: Thebaud, Thomas, et al.
Published: (2024)
by: Thebaud, Thomas, et al.
Published: (2024)
MATRA: Modeling the Attack Surface of Agentic AI Systems -- OpenClaw Case Study
by: Van hamme, Tim, et al.
Published: (2026)
by: Van hamme, Tim, et al.
Published: (2026)
Goal-Driven Risk Assessment for LLM-Powered Systems: A Healthcare Case Study
by: Nagaraja, Neha, et al.
Published: (2026)
by: Nagaraja, Neha, et al.
Published: (2026)
Measuring Safety Alignment Effects in Autonomous Security Agents
by: David, Isaac, et al.
Published: (2026)
by: David, Isaac, et al.
Published: (2026)
Targeting Alignment: Extracting Safety Classifiers of Aligned LLMs
by: Ferrand, Jean-Charles Noirot, et al.
Published: (2025)
by: Ferrand, Jean-Charles Noirot, et al.
Published: (2025)
Towards Compositional Generalization in LLMs for Smart Contract Security: A Case Study on Reentrancy Vulnerabilities
by: Zhou, Ying, et al.
Published: (2026)
by: Zhou, Ying, et al.
Published: (2026)
Preventing Adversarial AI Attacks Against Autonomous Situational Awareness: A Maritime Case Study
by: Walter, Mathew J., et al.
Published: (2025)
by: Walter, Mathew J., et al.
Published: (2025)
FreakOut-LLM: The Effect of Emotional Stimuli on Safety Alignment
by: Kuznetsov, Daniel, et al.
Published: (2026)
by: Kuznetsov, Daniel, et al.
Published: (2026)
VisuoAlign: Safety Alignment of LVLMs with Multimodal Tree Search
by: Li, MingSheng, et al.
Published: (2025)
by: Li, MingSheng, et al.
Published: (2025)
Case Study: Fine-tuning Small Language Models for Accurate and Private CWE Detection in Python Code
by: Bappy, Md. Azizul Hakim, et al.
Published: (2025)
by: Bappy, Md. Azizul Hakim, et al.
Published: (2025)
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications
by: David, Isaac, et al.
Published: (2026)
by: David, Isaac, et al.
Published: (2026)
Defensive Refusal Bias: How Safety Alignment Fails Cyber Defenders
by: Campbell, David, et al.
Published: (2026)
by: Campbell, David, et al.
Published: (2026)
Co-Evolutionary Multi-Modal Alignment via Structured Adversarial Evolution
by: Shi, Guoxin, et al.
Published: (2026)
by: Shi, Guoxin, et al.
Published: (2026)
When Alignment Isn't Enough: Response-Path Attacks on LLM Agents
by: Luo, Mingyu, et al.
Published: (2026)
by: Luo, Mingyu, et al.
Published: (2026)
Medical Multimodal Model Stealing Attacks via Adversarial Domain Alignment
by: Shen, Yaling, et al.
Published: (2025)
by: Shen, Yaling, et al.
Published: (2025)
Matching Ranks Over Probability Yields Truly Deep Safety Alignment
by: Vega, Jason, et al.
Published: (2025)
by: Vega, Jason, et al.
Published: (2025)
PRISM: Robust VLM Alignment with Principled Reasoning for Integrated Safety in Multimodality
by: Li, Nanxi, et al.
Published: (2025)
by: Li, Nanxi, et al.
Published: (2025)
Trojan Horses in Recruiting: A Red-Teaming Case Study on Indirect Prompt Injection in Standard vs. Reasoning Models
by: Wirth, Manuel
Published: (2026)
by: Wirth, Manuel
Published: (2026)
SafeThinker: Reasoning about Risk to Deepen Safety Beyond Shallow Alignment
by: Fang, Xianya, et al.
Published: (2026)
by: Fang, Xianya, et al.
Published: (2026)
Refusal Falls off a Cliff: How Safety Alignment Fails in Reasoning?
by: Yin, Qingyu, et al.
Published: (2025)
by: Yin, Qingyu, et al.
Published: (2025)
On the Impossibility of Separating Intelligence from Judgment: The Computational Intractability of Filtering for AI Alignment
by: Ball, Sarah, et al.
Published: (2025)
by: Ball, Sarah, et al.
Published: (2025)
Medoid Prototype Alignment for Cross-Plant Unknown Attack Detection in Industrial Control Systems
by: Wang, Luyao
Published: (2026)
by: Wang, Luyao
Published: (2026)
Safety Alignment Should Be Made More Than Just a Few Tokens Deep
by: Qi, Xiangyu, et al.
Published: (2024)
by: Qi, Xiangyu, et al.
Published: (2024)
MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming
by: Guo, Weiyang, et al.
Published: (2025)
by: Guo, Weiyang, et al.
Published: (2025)
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training
by: Du, Pengfei
Published: (2025)
by: Du, Pengfei
Published: (2025)
Towards Privacy-Preserving Large Language Model: Text-free Inference Through Alignment and Adaptation
by: Yoon, Jeongho, et al.
Published: (2026)
by: Yoon, Jeongho, et al.
Published: (2026)
Enhancing Decision-Making in Windows PE Malware Classification During Dataset Shifts with Uncertainty Estimation
by: Yumlembam, Rahul, et al.
Published: (2025)
by: Yumlembam, Rahul, et al.
Published: (2025)
Flow-based Detection of Botnets through Bio-inspired Optimisation of Machine Learning
by: Issac, Biju, et al.
Published: (2024)
by: Issac, Biju, et al.
Published: (2024)
Defending MoE LLMs against Harmful Fine-Tuning via Safety Routing Alignment
by: Kim, Jaehan, et al.
Published: (2025)
by: Kim, Jaehan, et al.
Published: (2025)
RAJ-PGA: Reasoning-Activated Jailbreak and Principle-Guided Alignment Framework for Large Reasoning Models
by: Chen, Jianhao, et al.
Published: (2025)
by: Chen, Jianhao, et al.
Published: (2025)
Similar Items
-
Evaluating whether AI models would sabotage AI safety research
by: Kirk, Robert, et al.
Published: (2026) -
Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents
by: Bazinska, Julia, et al.
Published: (2025) -
Fundamental Limitations in Pointwise Defences of LLM Finetuning APIs
by: Davies, Xander, et al.
Published: (2025) -
SeCodePLT: A Unified Platform for Evaluating the Security of Code GenAI
by: Nie, Yuzhou, et al.
Published: (2024) -
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
by: Che, Zora, et al.
Published: (2025)