SHAP zero Explains Biological Sequence Models with Near-zero Marginal Cost for Future Queries

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tsui, Darin, Musharaf, Aryan, Erginbas, Yigit Efe, Kang, Justin Singh, Aghazadeh, Amirali
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913852719890432
author Tsui, Darin
Musharaf, Aryan
Erginbas, Yigit Efe
Kang, Justin Singh
Aghazadeh, Amirali
author_facet Tsui, Darin
Musharaf, Aryan
Erginbas, Yigit Efe
Kang, Justin Singh
Aghazadeh, Amirali
contents The growing adoption of machine learning models for biological sequences has intensified the need for interpretable predictions, with Shapley values emerging as a theoretically grounded standard for model explanation. While effective for local explanations of individual input sequences, scaling Shapley-based interpretability to extract global biological insights requires evaluating thousands of sequences--incurring exponential computational cost per query. We introduce SHAP zero, a novel algorithm that amortizes the cost of Shapley value computation across large-scale biological datasets. After a one-time model sketching step, SHAP zero enables near-zero marginal cost for future queries by uncovering an underexplored connection between Shapley values, high-order feature interactions, and the sparse Fourier transform of the model. Applied to models of guide RNA efficacy, DNA repair outcomes, and protein fitness, SHAP zero explains predictions orders of magnitude faster than existing methods, recovering rich combinatorial interactions previously inaccessible at scale. This work opens the door to principled, efficient, and scalable interpretability for black-box sequence models in biology.
format Preprint
id arxiv_https___arxiv_org_abs_2410_19236
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SHAP zero Explains Biological Sequence Models with Near-zero Marginal Cost for Future Queries
Tsui, Darin
Musharaf, Aryan
Erginbas, Yigit Efe
Kang, Justin Singh
Aghazadeh, Amirali
Machine Learning
Computational Engineering, Finance, and Science
Genomics
Computation
The growing adoption of machine learning models for biological sequences has intensified the need for interpretable predictions, with Shapley values emerging as a theoretically grounded standard for model explanation. While effective for local explanations of individual input sequences, scaling Shapley-based interpretability to extract global biological insights requires evaluating thousands of sequences--incurring exponential computational cost per query. We introduce SHAP zero, a novel algorithm that amortizes the cost of Shapley value computation across large-scale biological datasets. After a one-time model sketching step, SHAP zero enables near-zero marginal cost for future queries by uncovering an underexplored connection between Shapley values, high-order feature interactions, and the sparse Fourier transform of the model. Applied to models of guide RNA efficacy, DNA repair outcomes, and protein fitness, SHAP zero explains predictions orders of magnitude faster than existing methods, recovering rich combinatorial interactions previously inaccessible at scale. This work opens the door to principled, efficient, and scalable interpretability for black-box sequence models in biology.
title SHAP zero Explains Biological Sequence Models with Near-zero Marginal Cost for Future Queries
topic Machine Learning
Computational Engineering, Finance, and Science
Genomics
Computation
url https://arxiv.org/abs/2410.19236