Antidistillation Sampling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Savani, Yash, Trockman, Asher, Feng, Zhili, Xu, Yixuan Even, Schwarzschild, Avi, Robey, Alexander, Finzi, Marc, Kolter, J. Zico
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909870383431680
author Savani, Yash
Trockman, Asher
Feng, Zhili
Xu, Yixuan Even
Schwarzschild, Avi
Robey, Alexander
Finzi, Marc
Kolter, J. Zico
author_facet Savani, Yash
Trockman, Asher
Feng, Zhili
Xu, Yixuan Even
Schwarzschild, Avi
Robey, Alexander
Finzi, Marc
Kolter, J. Zico
contents Frontier models that generate extended reasoning traces inadvertently produce rich token sequences that can facilitate model distillation. Recognizing this vulnerability, model owners may seek sampling strategies that limit the effectiveness of distillation without compromising model performance. Antidistillation sampling provides exactly this capability. By strategically modifying a model's next-token probability distribution, antidistillation sampling poisons reasoning traces, rendering them significantly less effective for distillation while preserving the model's practical utility. For further details, see https://antidistillation.com.
format Preprint
id arxiv_https___arxiv_org_abs_2504_13146
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Antidistillation Sampling
Savani, Yash
Trockman, Asher
Feng, Zhili
Xu, Yixuan Even
Schwarzschild, Avi
Robey, Alexander
Finzi, Marc
Kolter, J. Zico
Artificial Intelligence
Computation and Language
Frontier models that generate extended reasoning traces inadvertently produce rich token sequences that can facilitate model distillation. Recognizing this vulnerability, model owners may seek sampling strategies that limit the effectiveness of distillation without compromising model performance. Antidistillation sampling provides exactly this capability. By strategically modifying a model's next-token probability distribution, antidistillation sampling poisons reasoning traces, rendering them significantly less effective for distillation while preserving the model's practical utility. For further details, see https://antidistillation.com.
title Antidistillation Sampling
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2504.13146