AutoSciDACT: Automated Scientific Discovery through Contrastive Embedding and Hypothesis Testing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bright-Thonney, Samuel, Reissel, Christina, Grosso, Gaia, Woodward, Nathaniel, Govorkova, Katya, Novak, Andrzej, Park, Sang Eon, Moreno, Eric, Harris, Philip
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908783019556864
author Bright-Thonney, Samuel
Reissel, Christina
Grosso, Gaia
Woodward, Nathaniel
Govorkova, Katya
Novak, Andrzej
Park, Sang Eon
Moreno, Eric
Harris, Philip
author_facet Bright-Thonney, Samuel
Reissel, Christina
Grosso, Gaia
Woodward, Nathaniel
Govorkova, Katya
Novak, Andrzej
Park, Sang Eon
Moreno, Eric
Harris, Philip
contents Novelty detection in large scientific datasets faces two key challenges: the noisy and high-dimensional nature of experimental data, and the necessity of making statistically robust statements about any observed outliers. While there is a wealth of literature on anomaly detection via dimensionality reduction, most methods do not produce outputs compatible with quantifiable claims of scientific discovery. In this work we directly address these challenges, presenting the first step towards a unified pipeline for novelty detection adapted for the rigorous statistical demands of science. We introduce AutoSciDACT (Automated Scientific Discovery with Anomalous Contrastive Testing), a general-purpose pipeline for detecting novelty in scientific data. AutoSciDACT begins by creating expressive low-dimensional data representations using a contrastive pre-training, leveraging the abundance of high-quality simulated data in many scientific domains alongside expertise that can guide principled data augmentation strategies. These compact embeddings then enable an extremely sensitive machine learning-based two-sample test using the New Physics Learning Machine (NPLM) framework, which identifies and statistically quantifies deviations in observed data relative to a reference distribution (null hypothesis). We perform experiments across a range of astronomical, physical, biological, image, and synthetic datasets, demonstrating strong sensitivity to small injections of anomalous data across all domains.
format Preprint
id arxiv_https___arxiv_org_abs_2510_21935
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AutoSciDACT: Automated Scientific Discovery through Contrastive Embedding and Hypothesis Testing
Bright-Thonney, Samuel
Reissel, Christina
Grosso, Gaia
Woodward, Nathaniel
Govorkova, Katya
Novak, Andrzej
Park, Sang Eon
Moreno, Eric
Harris, Philip
Machine Learning
Artificial Intelligence
Novelty detection in large scientific datasets faces two key challenges: the noisy and high-dimensional nature of experimental data, and the necessity of making statistically robust statements about any observed outliers. While there is a wealth of literature on anomaly detection via dimensionality reduction, most methods do not produce outputs compatible with quantifiable claims of scientific discovery. In this work we directly address these challenges, presenting the first step towards a unified pipeline for novelty detection adapted for the rigorous statistical demands of science. We introduce AutoSciDACT (Automated Scientific Discovery with Anomalous Contrastive Testing), a general-purpose pipeline for detecting novelty in scientific data. AutoSciDACT begins by creating expressive low-dimensional data representations using a contrastive pre-training, leveraging the abundance of high-quality simulated data in many scientific domains alongside expertise that can guide principled data augmentation strategies. These compact embeddings then enable an extremely sensitive machine learning-based two-sample test using the New Physics Learning Machine (NPLM) framework, which identifies and statistically quantifies deviations in observed data relative to a reference distribution (null hypothesis). We perform experiments across a range of astronomical, physical, biological, image, and synthetic datasets, demonstrating strong sensitivity to small injections of anomalous data across all domains.
title AutoSciDACT: Automated Scientific Discovery through Contrastive Embedding and Hypothesis Testing
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2510.21935