ShadowGenes: Leveraging Recurring Patterns within Computational Graphs for Model Genealogy

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Schulz, Kasimir, Evans, Kieran
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917898467934208
author Schulz, Kasimir
Evans, Kieran
author_facet Schulz, Kasimir
Evans, Kieran
contents Machine learning model genealogy enables practitioners to determine which architectural family a neural network belongs to. In this paper, we introduce ShadowGenes, a novel, signature-based method for identifying a given model's architecture, type, and family. Our method involves building a computational graph of the model that is agnostic of its serialization format, then analyzing its internal operations to identify unique patterns, and finally building and refining signatures based on these. We highlight important workings of the underlying engine and demonstrate the technique used to construct a signature and scan a given model. This approach to model genealogy can be applied to model files without the need for additional external information. We test ShadowGenes on a labeled dataset of over 1,400 models and achieve a mean true positive rate of 97.49% and a precision score of 99.51%; which validates the technique as a practical method for model genealogy. This enables practitioners to understand the use cases of a given model, the internal computational process, and identify possible security risks, such as the potential for model backdooring.
format Preprint
id arxiv_https___arxiv_org_abs_2501_11830
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ShadowGenes: Leveraging Recurring Patterns within Computational Graphs for Model Genealogy
Schulz, Kasimir
Evans, Kieran
Machine Learning
Cryptography and Security
Machine learning model genealogy enables practitioners to determine which architectural family a neural network belongs to. In this paper, we introduce ShadowGenes, a novel, signature-based method for identifying a given model's architecture, type, and family. Our method involves building a computational graph of the model that is agnostic of its serialization format, then analyzing its internal operations to identify unique patterns, and finally building and refining signatures based on these. We highlight important workings of the underlying engine and demonstrate the technique used to construct a signature and scan a given model. This approach to model genealogy can be applied to model files without the need for additional external information. We test ShadowGenes on a labeled dataset of over 1,400 models and achieve a mean true positive rate of 97.49% and a precision score of 99.51%; which validates the technique as a practical method for model genealogy. This enables practitioners to understand the use cases of a given model, the internal computational process, and identify possible security risks, such as the potential for model backdooring.
title ShadowGenes: Leveraging Recurring Patterns within Computational Graphs for Model Genealogy
topic Machine Learning
Cryptography and Security
url https://arxiv.org/abs/2501.11830