Whole-Genome Phenotype Prediction with Machine Learning: Open Problems in Bacterial Genomics

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: James, Tamsin, Williamson, Ben, Tino, Peter, Wheeler, Nicole
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909488195305472
author James, Tamsin
Williamson, Ben
Tino, Peter
Wheeler, Nicole
author_facet James, Tamsin
Williamson, Ben
Tino, Peter
Wheeler, Nicole
contents How can we identify causal genetic mechanisms that govern bacterial traits? Initial efforts entrusting machine learning models to handle the task of predicting phenotype from genotype return high accuracy scores. However, attempts to extract any meaning from the predictive models are found to be corrupted by falsely identified "causal" features. Relying solely on pattern recognition and correlations is unreliable, significantly so in bacterial genomics settings where high-dimensionality and spurious associations are the norm. Though it is not yet clear whether we can overcome this hurdle, significant efforts are being made towards discovering potential high-risk bacterial genetic variants. In view of this, we set up open problems surrounding phenotype prediction from bacterial whole-genome datasets and extending those to learning causal effects, and discuss challenges that impact the reliability of a machine's decision-making when faced with datasets of this nature.
format Preprint
id arxiv_https___arxiv_org_abs_2502_07749
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Whole-Genome Phenotype Prediction with Machine Learning: Open Problems in Bacterial Genomics
James, Tamsin
Williamson, Ben
Tino, Peter
Wheeler, Nicole
Genomics
Machine Learning
How can we identify causal genetic mechanisms that govern bacterial traits? Initial efforts entrusting machine learning models to handle the task of predicting phenotype from genotype return high accuracy scores. However, attempts to extract any meaning from the predictive models are found to be corrupted by falsely identified "causal" features. Relying solely on pattern recognition and correlations is unreliable, significantly so in bacterial genomics settings where high-dimensionality and spurious associations are the norm. Though it is not yet clear whether we can overcome this hurdle, significant efforts are being made towards discovering potential high-risk bacterial genetic variants. In view of this, we set up open problems surrounding phenotype prediction from bacterial whole-genome datasets and extending those to learning causal effects, and discuss challenges that impact the reliability of a machine's decision-making when faced with datasets of this nature.
title Whole-Genome Phenotype Prediction with Machine Learning: Open Problems in Bacterial Genomics
topic Genomics
Machine Learning
url https://arxiv.org/abs/2502.07749