Leave It to the Experts: Detecting Knowledge Distillation via MoE Expert Signatures

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Pingzhi, Huang, Morris Yu-Chao, Tan, Zhen, Song, Qingquan, Peng, Jie, Zou, Kai, Cheng, Yu, Xu, Kaidi, Chen, Tianlong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911220358971392
author Li, Pingzhi
Huang, Morris Yu-Chao
Tan, Zhen
Song, Qingquan
Peng, Jie
Zou, Kai
Cheng, Yu
Xu, Kaidi
Chen, Tianlong
author_facet Li, Pingzhi
Huang, Morris Yu-Chao
Tan, Zhen
Song, Qingquan
Peng, Jie
Zou, Kai
Cheng, Yu
Xu, Kaidi
Chen, Tianlong
contents Knowledge Distillation (KD) accelerates training of large language models (LLMs) but poses intellectual property protection and LLM diversity risks. Existing KD detection methods based on self-identity or output similarity can be easily evaded through prompt engineering. We present a KD detection framework effective in both white-box and black-box settings by exploiting an overlooked signal: the transfer of MoE "structural habits", especially internal routing patterns. Our approach analyzes how different experts specialize and collaborate across various inputs, creating distinctive fingerprints that persist through the distillation process. To extend beyond the white-box setup and MoE architectures, we further propose Shadow-MoE, a black-box method that constructs proxy MoE representations via auxiliary distillation to compare these patterns between arbitrary model pairs. We establish a comprehensive, reproducible benchmark that offers diverse distilled checkpoints and an extensible framework to facilitate future research. Extensive experiments demonstrate >94% detection accuracy across various scenarios and strong robustness to prompt-based evasion, outperforming existing baselines while highlighting the structural habits transfer in LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2510_16968
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Leave It to the Experts: Detecting Knowledge Distillation via MoE Expert Signatures
Li, Pingzhi
Huang, Morris Yu-Chao
Tan, Zhen
Song, Qingquan
Peng, Jie
Zou, Kai
Cheng, Yu
Xu, Kaidi
Chen, Tianlong
Machine Learning
Artificial Intelligence
Computation and Language
Knowledge Distillation (KD) accelerates training of large language models (LLMs) but poses intellectual property protection and LLM diversity risks. Existing KD detection methods based on self-identity or output similarity can be easily evaded through prompt engineering. We present a KD detection framework effective in both white-box and black-box settings by exploiting an overlooked signal: the transfer of MoE "structural habits", especially internal routing patterns. Our approach analyzes how different experts specialize and collaborate across various inputs, creating distinctive fingerprints that persist through the distillation process. To extend beyond the white-box setup and MoE architectures, we further propose Shadow-MoE, a black-box method that constructs proxy MoE representations via auxiliary distillation to compare these patterns between arbitrary model pairs. We establish a comprehensive, reproducible benchmark that offers diverse distilled checkpoints and an extensible framework to facilitate future research. Extensive experiments demonstrate >94% detection accuracy across various scenarios and strong robustness to prompt-based evasion, outperforming existing baselines while highlighting the structural habits transfer in LLMs.
title Leave It to the Experts: Detecting Knowledge Distillation via MoE Expert Signatures
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2510.16968