Acceptance Cards:A Four-Diagnostic Standard for Safe Fine-Tuning Defense Claims

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Konrad, Phongsakon Mark, Tanyel, Toygar, Ayvaz, Serkan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913112663261184
author Konrad, Phongsakon Mark
Tanyel, Toygar
Ayvaz, Serkan
author_facet Konrad, Phongsakon Mark
Tanyel, Toygar
Ayvaz, Serkan
contents Safe fine-tuning defenses are often endorsed on the basis of a held-out gap reduction, but the same reduction can come from sampling noise, subject artifacts, capability loss, or a mechanism that does not transfer. We introduce Acceptance Cards: an evaluation protocol, a documentation object, an executable audit package, and a claim-specific evidential standard for safe fine-tuning defense claims. The protocol checks statistical reliability, fresh semantic generalization, mechanism alignment, and cross-task transfer before treating a gap reduction as a full-card pass. Re-scored under this installed-gap protocol, SafeLoRA fails the full-card pass on Gemma-2-2B-it: under strict mechanism-class coding it fails all four diagnostics, and under a permissive shrinkage relabel it still fails three of four. This is a narrow installed-gap audit on one model family, not a global judgment of SafeLoRA's effectiveness. In a 46-cell audit, no cell satisfies the strict conjunction. The closest family is a near miss that passes reliability and mechanism checks where the required data are available, but fails the fresh-subject threshold, lacks a strict transfer pass, and carries a measurable deployment-accuracy cost.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10575
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Acceptance Cards:A Four-Diagnostic Standard for Safe Fine-Tuning Defense Claims
Konrad, Phongsakon Mark
Tanyel, Toygar
Ayvaz, Serkan
Cryptography and Security
Artificial Intelligence
Machine Learning
Safe fine-tuning defenses are often endorsed on the basis of a held-out gap reduction, but the same reduction can come from sampling noise, subject artifacts, capability loss, or a mechanism that does not transfer. We introduce Acceptance Cards: an evaluation protocol, a documentation object, an executable audit package, and a claim-specific evidential standard for safe fine-tuning defense claims. The protocol checks statistical reliability, fresh semantic generalization, mechanism alignment, and cross-task transfer before treating a gap reduction as a full-card pass. Re-scored under this installed-gap protocol, SafeLoRA fails the full-card pass on Gemma-2-2B-it: under strict mechanism-class coding it fails all four diagnostics, and under a permissive shrinkage relabel it still fails three of four. This is a narrow installed-gap audit on one model family, not a global judgment of SafeLoRA's effectiveness. In a 46-cell audit, no cell satisfies the strict conjunction. The closest family is a near miss that passes reliability and mechanism checks where the required data are available, but fails the fresh-subject threshold, lacks a strict transfer pass, and carries a measurable deployment-accuracy cost.
title Acceptance Cards:A Four-Diagnostic Standard for Safe Fine-Tuning Defense Claims
topic Cryptography and Security
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2605.10575