An Example Safety Case for Safeguards Against Misuse
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Clymer, Joshua, Weinbaum, Jonah, Kirk, Robert, Mai, Kimberly, Zhang, Selena, Davies, Xander |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs
von: O'Brien, Kyle, et al.
Veröffentlicht: (2025)
von: O'Brien, Kyle, et al.
Veröffentlicht: (2025)
Safety Cases: How to Justify the Safety of Advanced AI Systems
von: Clymer, Joshua, et al.
Veröffentlicht: (2024)
von: Clymer, Joshua, et al.
Veröffentlicht: (2024)
Existing Large Language Model Unlearning Evaluations Are Inconclusive
von: Feng, Zhili, et al.
Veröffentlicht: (2025)
von: Feng, Zhili, et al.
Veröffentlicht: (2025)
TERD: A Unified Framework for Safeguarding Diffusion Models Against Backdoors
von: Mo, Yichuan, et al.
Veröffentlicht: (2024)
von: Mo, Yichuan, et al.
Veröffentlicht: (2024)
Adversaries Can Misuse Combinations of Safe Models
von: Jones, Erik, et al.
Veröffentlicht: (2024)
von: Jones, Erik, et al.
Veröffentlicht: (2024)
STACK: Adversarial Attacks on LLM Safeguard Pipelines
von: McKenzie, Ian R., et al.
Veröffentlicht: (2025)
von: McKenzie, Ian R., et al.
Veröffentlicht: (2025)
PFGuard: A Generative Framework with Privacy and Fairness Safeguards
von: Kim, Soyeon, et al.
Veröffentlicht: (2024)
von: Kim, Soyeon, et al.
Veröffentlicht: (2024)
UK AISI Alignment Evaluation Case-Study
von: Souly, Alexandra, et al.
Veröffentlicht: (2026)
von: Souly, Alexandra, et al.
Veröffentlicht: (2026)
NExT-Guard: Training-Free Streaming Safeguard without Token-Level Labels
von: Fang, Junfeng, et al.
Veröffentlicht: (2026)
von: Fang, Junfeng, et al.
Veröffentlicht: (2026)
ML-On-Rails: Safeguarding Machine Learning Models in Software Systems A Case Study
von: Abdelkader, Hala, et al.
Veröffentlicht: (2024)
von: Abdelkader, Hala, et al.
Veröffentlicht: (2024)
Efficient Safety Retrofitting Against Jailbreaking for LLMs
von: Garcia-Gasulla, Dario, et al.
Veröffentlicht: (2025)
von: Garcia-Gasulla, Dario, et al.
Veröffentlicht: (2025)
CORA: Conformal Risk-Controlled Agents for Safeguarded Mobile GUI Automation
von: Feng, Yushi, et al.
Veröffentlicht: (2026)
von: Feng, Yushi, et al.
Veröffentlicht: (2026)
Safeguarding LLM Fine-tuning via Push-Pull Distributional Alignment
von: Wang, Haozhong, et al.
Veröffentlicht: (2026)
von: Wang, Haozhong, et al.
Veröffentlicht: (2026)
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs
von: Fonseca, Joao, et al.
Veröffentlicht: (2025)
von: Fonseca, Joao, et al.
Veröffentlicht: (2025)
Alert-ME: An Explainability-Driven Defense Against Adversarial Examples in Transformer-Based Text Classification
von: Sabir, Bushra, et al.
Veröffentlicht: (2023)
von: Sabir, Bushra, et al.
Veröffentlicht: (2023)
Diet-ODIN: A Novel Framework for Opioid Misuse Detection with Interpretable Dietary Patterns
von: Zhang, Zheyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Zheyuan, et al.
Veröffentlicht: (2024)
SABER: Small Actions, Big Errors -- Safeguarding Mutating Steps in LLM Agents
von: Cuadron, Alejandro, et al.
Veröffentlicht: (2025)
von: Cuadron, Alejandro, et al.
Veröffentlicht: (2025)
Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents
von: Bazinska, Julia, et al.
Veröffentlicht: (2025)
von: Bazinska, Julia, et al.
Veröffentlicht: (2025)
GuardReasoner: Towards Reasoning-based LLM Safeguards
von: Liu, Yue, et al.
Veröffentlicht: (2025)
von: Liu, Yue, et al.
Veröffentlicht: (2025)
Towards a Novel Perspective on Adversarial Examples Driven by Frequency
von: Zhang, Zhun, et al.
Veröffentlicht: (2024)
von: Zhang, Zhun, et al.
Veröffentlicht: (2024)
PropGuard: Safeguarding LLM-MAS via Propagation-Aware Exploration and Remediation
von: Yan, Bingyu, et al.
Veröffentlicht: (2026)
von: Yan, Bingyu, et al.
Veröffentlicht: (2026)
On Prompt-Driven Safeguarding for Large Language Models
von: Zheng, Chujie, et al.
Veröffentlicht: (2024)
von: Zheng, Chujie, et al.
Veröffentlicht: (2024)
Tamper-Resistant Safeguards for Open-Weight LLMs
von: Tamirisa, Rishub, et al.
Veröffentlicht: (2024)
von: Tamirisa, Rishub, et al.
Veröffentlicht: (2024)
Debiasing Machine Unlearning with Counterfactual Examples
von: Chen, Ziheng, et al.
Veröffentlicht: (2024)
von: Chen, Ziheng, et al.
Veröffentlicht: (2024)
Kelp: A Streaming Safeguard for Large Models via Latent Dynamics-Guided Risk Detection
von: Li, Xiaodan, et al.
Veröffentlicht: (2025)
von: Li, Xiaodan, et al.
Veröffentlicht: (2025)
Difficult Examples Hurt Unsupervised Contrastive Learning: A Theoretical Perspective
von: Zhang, Yi-Ge, et al.
Veröffentlicht: (2025)
von: Zhang, Yi-Ge, et al.
Veröffentlicht: (2025)
Reward Model Overoptimisation in Iterated RLHF
von: Wolf, Lorenz, et al.
Veröffentlicht: (2025)
von: Wolf, Lorenz, et al.
Veröffentlicht: (2025)
How Do Large Language Monkeys Get Their Power (Laws)?
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025)
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025)
Predicting Performance of Symbolic and Prompt Programs with Examples
von: Zheng, Chengqi, et al.
Veröffentlicht: (2026)
von: Zheng, Chengqi, et al.
Veröffentlicht: (2026)
TabGen-ICL: Residual-Aware In-Context Example Selection for Tabular Data Generation
von: Fang, Liancheng, et al.
Veröffentlicht: (2025)
von: Fang, Liancheng, et al.
Veröffentlicht: (2025)
A Concrete Roadmap towards Safety Cases based on Chain-of-Thought Monitoring
von: Schulz, Julian
Veröffentlicht: (2025)
von: Schulz, Julian
Veröffentlicht: (2025)
Conformal Safety Monitoring for Flight Testing: A Case Study in Data-Driven Safety Learning
von: Feldman, Aaron O., et al.
Veröffentlicht: (2025)
von: Feldman, Aaron O., et al.
Veröffentlicht: (2025)
Analyzing the Impact of Adversarial Examples on Explainable Machine Learning
von: Devabhakthini, Prathyusha, et al.
Veröffentlicht: (2023)
von: Devabhakthini, Prathyusha, et al.
Veröffentlicht: (2023)
Safety Modulation: Enhancing Safety in Reinforcement Learning through Cost-Modulated Rewards
von: Zhang, Hanping, et al.
Veröffentlicht: (2025)
von: Zhang, Hanping, et al.
Veröffentlicht: (2025)
Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards
von: Pandey, Punya Syon, et al.
Veröffentlicht: (2025)
von: Pandey, Punya Syon, et al.
Veröffentlicht: (2025)
Safeguarding Autonomy: a Focus on Machine Learning Decision Systems
von: Subías-Beltrán, Paula, et al.
Veröffentlicht: (2025)
von: Subías-Beltrán, Paula, et al.
Veröffentlicht: (2025)
MemLoss: Enhancing Adversarial Training with Recycling Adversarial Examples
von: Mahdi, Soroush, et al.
Veröffentlicht: (2025)
von: Mahdi, Soroush, et al.
Veröffentlicht: (2025)
Towards Interpretable Adversarial Examples via Sparse Adversarial Attack
von: Lin, Fudong, et al.
Veröffentlicht: (2025)
von: Lin, Fudong, et al.
Veröffentlicht: (2025)
Obtaining Example-Based Explanations from Deep Neural Networks
von: Dong, Genghua, et al.
Veröffentlicht: (2025)
von: Dong, Genghua, et al.
Veröffentlicht: (2025)
Addressing the Abstraction and Reasoning Corpus via Procedural Example Generation
von: Hodel, Michael
Veröffentlicht: (2024)
von: Hodel, Michael
Veröffentlicht: (2024)
Ähnliche Einträge
-
Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs
von: O'Brien, Kyle, et al.
Veröffentlicht: (2025) -
Safety Cases: How to Justify the Safety of Advanced AI Systems
von: Clymer, Joshua, et al.
Veröffentlicht: (2024) -
Existing Large Language Model Unlearning Evaluations Are Inconclusive
von: Feng, Zhili, et al.
Veröffentlicht: (2025) -
TERD: A Unified Framework for Safeguarding Diffusion Models Against Backdoors
von: Mo, Yichuan, et al.
Veröffentlicht: (2024) -
Adversaries Can Misuse Combinations of Safe Models
von: Jones, Erik, et al.
Veröffentlicht: (2024)