Interpreting GFlowNets for Drug Discovery: Extracting Actionable Insights for Medicinal Chemistry

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: S, Amirtha Varshini A, Ranasinghe, Duminda S., Tam, Hok Hei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912726915219456
author S, Amirtha Varshini A
Ranasinghe, Duminda S.
Tam, Hok Hei
author_facet S, Amirtha Varshini A
Ranasinghe, Duminda S.
Tam, Hok Hei
contents Generative Flow Networks, or GFlowNets, offer a promising framework for molecular design, but their internal decision policies remain opaque. This limits adoption in drug discovery, where chemists require clear and interpretable rationales for proposed structures. We present an interpretability framework for SynFlowNet, a GFlowNet trained on documented chemical reactions and purchasable starting materials that generates both molecules and the synthetic routes that produce them. Our approach integrates three complementary components. Gradient based saliency combined with counterfactual perturbations identifies which atomic environments influence reward and how structural edits change molecular outcomes. Sparse autoencoders reveal axis aligned latent factors that correspond to physicochemical properties such as polarity, lipophilicity, and molecular size. Motif probes show that functional groups including aromatic rings and halogens are explicitly encoded and linearly decodable from the internal embeddings. Together, these results expose the chemical logic inside SynFlowNet and provide actionable and mechanistic insight that supports transparent and controllable molecular design.
format Preprint
id arxiv_https___arxiv_org_abs_2511_19264
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Interpreting GFlowNets for Drug Discovery: Extracting Actionable Insights for Medicinal Chemistry
S, Amirtha Varshini A
Ranasinghe, Duminda S.
Tam, Hok Hei
Machine Learning
Artificial Intelligence
Biomolecules
Generative Flow Networks, or GFlowNets, offer a promising framework for molecular design, but their internal decision policies remain opaque. This limits adoption in drug discovery, where chemists require clear and interpretable rationales for proposed structures. We present an interpretability framework for SynFlowNet, a GFlowNet trained on documented chemical reactions and purchasable starting materials that generates both molecules and the synthetic routes that produce them. Our approach integrates three complementary components. Gradient based saliency combined with counterfactual perturbations identifies which atomic environments influence reward and how structural edits change molecular outcomes. Sparse autoencoders reveal axis aligned latent factors that correspond to physicochemical properties such as polarity, lipophilicity, and molecular size. Motif probes show that functional groups including aromatic rings and halogens are explicitly encoded and linearly decodable from the internal embeddings. Together, these results expose the chemical logic inside SynFlowNet and provide actionable and mechanistic insight that supports transparent and controllable molecular design.
title Interpreting GFlowNets for Drug Discovery: Extracting Actionable Insights for Medicinal Chemistry
topic Machine Learning
Artificial Intelligence
Biomolecules
url https://arxiv.org/abs/2511.19264