Eliciting Harmful Capabilities by Fine-Tuning On Safeguarded Outputs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kaunismaa, Jackson, Griffin, Avery, Hughes, John, Knight, Christina Q., Sharma, Mrinank, Jones, Erik
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908775627096064
author Kaunismaa, Jackson
Griffin, Avery
Hughes, John
Knight, Christina Q.
Sharma, Mrinank
Jones, Erik
author_facet Kaunismaa, Jackson
Griffin, Avery
Hughes, John
Knight, Christina Q.
Sharma, Mrinank
Jones, Erik
contents Model developers implement safeguards in frontier models to prevent misuse, for example, by employing classifiers to filter dangerous outputs. In this work, we demonstrate that even robustly safeguarded models can be used to elicit harmful capabilities in open-source models through elicitation attacks. Our elicitation attacks consist of three stages: (i) constructing prompts in adjacent domains to a target harmful task that do not request dangerous information; (ii) obtaining responses to these prompts from safeguarded frontier models; (iii) fine-tuning open-source models on these prompt-output pairs. Since the requested prompts cannot be used to directly cause harm, they are not refused by frontier model safeguards. We evaluate these elicitation attacks within the domain of hazardous chemical synthesis and processing, and demonstrate that our attacks recover approximately 40% of the capability gap between the base open-source model and an unrestricted frontier model. We then show that the efficacy of elicitation attacks scales with the capability of the frontier model and the amount of generated fine-tuning data. Our work demonstrates the challenge of mitigating ecosystem level risks with output-level safeguards.
format Preprint
id arxiv_https___arxiv_org_abs_2601_13528
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Eliciting Harmful Capabilities by Fine-Tuning On Safeguarded Outputs
Kaunismaa, Jackson
Griffin, Avery
Hughes, John
Knight, Christina Q.
Sharma, Mrinank
Jones, Erik
Cryptography and Security
Artificial Intelligence
Computation and Language
Machine Learning
Software Engineering
Model developers implement safeguards in frontier models to prevent misuse, for example, by employing classifiers to filter dangerous outputs. In this work, we demonstrate that even robustly safeguarded models can be used to elicit harmful capabilities in open-source models through elicitation attacks. Our elicitation attacks consist of three stages: (i) constructing prompts in adjacent domains to a target harmful task that do not request dangerous information; (ii) obtaining responses to these prompts from safeguarded frontier models; (iii) fine-tuning open-source models on these prompt-output pairs. Since the requested prompts cannot be used to directly cause harm, they are not refused by frontier model safeguards. We evaluate these elicitation attacks within the domain of hazardous chemical synthesis and processing, and demonstrate that our attacks recover approximately 40% of the capability gap between the base open-source model and an unrestricted frontier model. We then show that the efficacy of elicitation attacks scales with the capability of the frontier model and the amount of generated fine-tuning data. Our work demonstrates the challenge of mitigating ecosystem level risks with output-level safeguards.
title Eliciting Harmful Capabilities by Fine-Tuning On Safeguarded Outputs
topic Cryptography and Security
Artificial Intelligence
Computation and Language
Machine Learning
Software Engineering
url https://arxiv.org/abs/2601.13528