Semantically Rich Local Dataset Generation for Explainable AI in Genomics

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Barbosa, Pedro, Savisaar, Rosina, Fonseca, Alcides
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914874192297984
author Barbosa, Pedro
Savisaar, Rosina
Fonseca, Alcides
author_facet Barbosa, Pedro
Savisaar, Rosina
Fonseca, Alcides
contents Black box deep learning models trained on genomic sequences excel at predicting the outcomes of different gene regulatory mechanisms. Therefore, interpreting these models may provide novel insights into the underlying biology, supporting downstream biomedical applications. Due to their complexity, interpretable surrogate models can only be built for local explanations (e.g., a single instance). However, accomplishing this requires generating a dataset in the neighborhood of the input, which must maintain syntactic similarity to the original data while introducing semantic variability in the model's predictions. This task is challenging due to the complex sequence-to-function relationship of DNA. We propose using Genetic Programming to generate datasets by evolving perturbations in sequences that contribute to their semantic diversity. Our custom, domain-guided individual representation effectively constrains syntactic similarity, and we provide two alternative fitness functions that promote diversity with no computational effort. Applied to the RNA splicing domain, our approach quickly achieves good diversity and significantly outperforms a random baseline in exploring the search space, as shown by our proof-of-concept, short RNA sequence. Furthermore, we assess its generalizability and demonstrate scalability to larger sequences, resulting in a ~30% improvement over the baseline.
format Preprint
id arxiv_https___arxiv_org_abs_2407_02984
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Semantically Rich Local Dataset Generation for Explainable AI in Genomics
Barbosa, Pedro
Savisaar, Rosina
Fonseca, Alcides
Machine Learning
Neural and Evolutionary Computing
Genomics
Black box deep learning models trained on genomic sequences excel at predicting the outcomes of different gene regulatory mechanisms. Therefore, interpreting these models may provide novel insights into the underlying biology, supporting downstream biomedical applications. Due to their complexity, interpretable surrogate models can only be built for local explanations (e.g., a single instance). However, accomplishing this requires generating a dataset in the neighborhood of the input, which must maintain syntactic similarity to the original data while introducing semantic variability in the model's predictions. This task is challenging due to the complex sequence-to-function relationship of DNA. We propose using Genetic Programming to generate datasets by evolving perturbations in sequences that contribute to their semantic diversity. Our custom, domain-guided individual representation effectively constrains syntactic similarity, and we provide two alternative fitness functions that promote diversity with no computational effort. Applied to the RNA splicing domain, our approach quickly achieves good diversity and significantly outperforms a random baseline in exploring the search space, as shown by our proof-of-concept, short RNA sequence. Furthermore, we assess its generalizability and demonstrate scalability to larger sequences, resulting in a ~30% improvement over the baseline.
title Semantically Rich Local Dataset Generation for Explainable AI in Genomics
topic Machine Learning
Neural and Evolutionary Computing
Genomics
url https://arxiv.org/abs/2407.02984