On Evaluating the Durability of Safeguards for Open-Weight LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qi, Xiangyu, Wei, Boyi, Carlini, Nicholas, Huang, Yangsibo, Xie, Tinghao, He, Luxi, Jagielski, Matthew, Nasr, Milad, Mittal, Prateek, Henderson, Peter |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Remote Timing Attacks on Efficient Language Model Inference
von: Carlini, Nicholas, et al.
Veröffentlicht: (2024)
von: Carlini, Nicholas, et al.
Veröffentlicht: (2024)
LLMs unlock new paths to monetizing exploits
von: Carlini, Nicholas, et al.
Veröffentlicht: (2025)
von: Carlini, Nicholas, et al.
Veröffentlicht: (2025)
Auditing Private Prediction
von: Chadha, Karan, et al.
Veröffentlicht: (2024)
von: Chadha, Karan, et al.
Veröffentlicht: (2024)
Privacy Side Channels in Machine Learning Systems
von: Debenedetti, Edoardo, et al.
Veröffentlicht: (2023)
von: Debenedetti, Edoardo, et al.
Veröffentlicht: (2023)
Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards
von: Huang, Yangsibo, et al.
Veröffentlicht: (2025)
von: Huang, Yangsibo, et al.
Veröffentlicht: (2025)
Private Fine-tuning of Large Language Models with Zeroth-order Optimization
von: Tang, Xinyu, et al.
Veröffentlicht: (2024)
von: Tang, Xinyu, et al.
Veröffentlicht: (2024)
Cascading Adversarial Bias from Injection to Distillation in Language Models
von: Chaudhari, Harsh, et al.
Veröffentlicht: (2025)
von: Chaudhari, Harsh, et al.
Veröffentlicht: (2025)
AI Risk Management Should Incorporate Both Safety and Security
von: Qi, Xiangyu, et al.
Veröffentlicht: (2024)
von: Qi, Xiangyu, et al.
Veröffentlicht: (2024)
Privacy Auditing of Large Language Models
von: Panda, Ashwinee, et al.
Veröffentlicht: (2025)
von: Panda, Ashwinee, et al.
Veröffentlicht: (2025)
AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
von: Carlini, Nicholas, et al.
Veröffentlicht: (2025)
von: Carlini, Nicholas, et al.
Veröffentlicht: (2025)
Stealing Part of a Production Language Model
von: Carlini, Nicholas, et al.
Veröffentlicht: (2024)
von: Carlini, Nicholas, et al.
Veröffentlicht: (2024)
Synthetic Query Generation for Privacy-Preserving Deep Retrieval Systems using Differentially Private Language Models
von: Carranza, Aldo Gael, et al.
Veröffentlicht: (2023)
von: Carranza, Aldo Gael, et al.
Veröffentlicht: (2023)
Thought-Transfer: Indirect Targeted Poisoning Attacks on Chain-of-Thought Reasoning Models
von: Chaudhari, Harsh, et al.
Veröffentlicht: (2026)
von: Chaudhari, Harsh, et al.
Veröffentlicht: (2026)
Query-Based Adversarial Prompt Generation
von: Hayase, Jonathan, et al.
Veröffentlicht: (2024)
von: Hayase, Jonathan, et al.
Veröffentlicht: (2024)
Safety Alignment Should Be Made More Than Just a Few Tokens Deep
von: Qi, Xiangyu, et al.
Veröffentlicht: (2024)
von: Qi, Xiangyu, et al.
Veröffentlicht: (2024)
Unlearn and Burn: Adversarial Machine Unlearning Requests Destroy Model Accuracy
von: Huang, Yangsibo, et al.
Veröffentlicht: (2024)
von: Huang, Yangsibo, et al.
Veröffentlicht: (2024)
Best Practices for Biorisk Evaluations on Open-Weight Bio-Foundation Models
von: Wei, Boyi, et al.
Veröffentlicht: (2025)
von: Wei, Boyi, et al.
Veröffentlicht: (2025)
Evaluating the Robustness of a Production Malware Detection System to Transferable Adversarial Attacks
von: Nasr, Milad, et al.
Veröffentlicht: (2025)
von: Nasr, Milad, et al.
Veröffentlicht: (2025)
An Adversarial Perspective on Machine Unlearning for AI Safety
von: Łucki, Jakub, et al.
Veröffentlicht: (2024)
von: Łucki, Jakub, et al.
Veröffentlicht: (2024)
Are aligned neural networks adversarially aligned?
von: Carlini, Nicholas, et al.
Veröffentlicht: (2023)
von: Carlini, Nicholas, et al.
Veröffentlicht: (2023)
Cutting through buggy adversarial example defenses: fixing 1 line of code breaks Sabre
von: Carlini, Nicholas
Veröffentlicht: (2024)
von: Carlini, Nicholas
Veröffentlicht: (2024)
PromptKeeper: Safeguarding System Prompts for LLMs
von: Jiang, Zhifeng, et al.
Veröffentlicht: (2024)
von: Jiang, Zhifeng, et al.
Veröffentlicht: (2024)
What is in Your Safe Data? Identifying Benign Data that Breaks Safety
von: He, Luxi, et al.
Veröffentlicht: (2024)
von: He, Luxi, et al.
Veröffentlicht: (2024)
Poisoning Web-Scale Training Datasets is Practical
von: Carlini, Nicholas, et al.
Veröffentlicht: (2023)
von: Carlini, Nicholas, et al.
Veröffentlicht: (2023)
The Last Iterate Advantage: Empirical Auditing and Principled Heuristic Analysis of Differentially Private SGD
von: Steinke, Thomas, et al.
Veröffentlicht: (2024)
von: Steinke, Thomas, et al.
Veröffentlicht: (2024)
IF-GUIDE: Influence Function-Guided Detoxification of LLMs
von: Coalson, Zachary, et al.
Veröffentlicht: (2025)
von: Coalson, Zachary, et al.
Veröffentlicht: (2025)
Hush! Protecting Secrets During Model Training: An Indistinguishability Approach
von: Ganesh, Arun, et al.
Veröffentlicht: (2025)
von: Ganesh, Arun, et al.
Veröffentlicht: (2025)
Phantom: General Backdoor Attacks on Retrieval Augmented Language Generation
von: Chaudhari, Harsh, et al.
Veröffentlicht: (2024)
von: Chaudhari, Harsh, et al.
Veröffentlicht: (2024)
Do Not Merge My Model! Safeguarding Open-Source LLMs Against Unauthorized Model Merging
von: Li, Qinfeng, et al.
Veröffentlicht: (2025)
von: Li, Qinfeng, et al.
Veröffentlicht: (2025)
Safeguarding LLMs Against Misuse and AI-Driven Malware Using Steganographic Canaries
von: Raz, Md, et al.
Veröffentlicht: (2026)
von: Raz, Md, et al.
Veröffentlicht: (2026)
Adversarial ML Problems Are Getting Harder to Solve and to Evaluate
von: Rando, Javier, et al.
Veröffentlicht: (2025)
von: Rando, Javier, et al.
Veröffentlicht: (2025)
Covert Attacks on Machine Learning Training in Passively Secure MPC
von: Jagielski, Matthew, et al.
Veröffentlicht: (2025)
von: Jagielski, Matthew, et al.
Veröffentlicht: (2025)
Position: Towards Resilience Against Adversarial Examples
von: Dai, Sihui, et al.
Veröffentlicht: (2024)
von: Dai, Sihui, et al.
Veröffentlicht: (2024)
Adversarial Perturbations Cannot Reliably Protect Artists From Generative AI
von: Hönig, Robert, et al.
Veröffentlicht: (2024)
von: Hönig, Robert, et al.
Veröffentlicht: (2024)
Re-Triggering Safeguards within LLMs for Jailbreak Detection
von: Lin, Zheng, et al.
Veröffentlicht: (2026)
von: Lin, Zheng, et al.
Veröffentlicht: (2026)
Context manipulation attacks : Web agents are susceptible to corrupted memory
von: Patlan, Atharv Singh, et al.
Veröffentlicht: (2025)
von: Patlan, Atharv Singh, et al.
Veröffentlicht: (2025)
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
von: Wang, Zhun, et al.
Veröffentlicht: (2026)
von: Wang, Zhun, et al.
Veröffentlicht: (2026)
CellSecInspector: Safeguarding Cellular Networks via Automated Security Analysis on Specifications
von: Xie, Ke, et al.
Veröffentlicht: (2025)
von: Xie, Ke, et al.
Veröffentlicht: (2025)
Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
von: Xiong, Chen, et al.
Veröffentlicht: (2024)
von: Xiong, Chen, et al.
Veröffentlicht: (2024)
Detecting Adversarial Fine-tuning with Auditing Agents
von: Egler, Sarah, et al.
Veröffentlicht: (2025)
von: Egler, Sarah, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Remote Timing Attacks on Efficient Language Model Inference
von: Carlini, Nicholas, et al.
Veröffentlicht: (2024) -
LLMs unlock new paths to monetizing exploits
von: Carlini, Nicholas, et al.
Veröffentlicht: (2025) -
Auditing Private Prediction
von: Chadha, Karan, et al.
Veröffentlicht: (2024) -
Privacy Side Channels in Machine Learning Systems
von: Debenedetti, Edoardo, et al.
Veröffentlicht: (2023) -
Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards
von: Huang, Yangsibo, et al.
Veröffentlicht: (2025)