CAP: Commutative Algebra Prediction of Protein-Nucleic Acid Binding Affinities

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zia, Mushal, Suwayyid, Faisal, Hozumi, Yuta, Wee, JunJie, Feng, Hongsong, Wei, Guo-Wei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911231279890432
author Zia, Mushal
Suwayyid, Faisal
Hozumi, Yuta
Wee, JunJie
Feng, Hongsong
Wei, Guo-Wei
author_facet Zia, Mushal
Suwayyid, Faisal
Hozumi, Yuta
Wee, JunJie
Feng, Hongsong
Wei, Guo-Wei
contents An accurate prediction of protein-nucleic acid binding affinity is vital for deciphering genomic processes, yet existing approaches often struggle in reconciling high accuracy with interpretability and computational efficiency. In this study, we introduce commutative algebra prediction (CAP), which couples persistent Stanley-Reisner theory with advanced sequence embedding for predicting protein-nucleic acid binding affinities. CAP encodes proteins through transformer-learned embeddings that retain long-range evolutionary context and represents DNA and RNA with $\textit{k}$-mer algebra embeddings derived from persistent facet ideals, which capture fine-scale nucleotide geometry. We demonstrate that CAP surpasses the SVSBI protein-nucleic acid benchmark and, in a further test, maintains reasonable performance on newly curated protein-RNA and protein-nucleic acid datasets. Leveraging only primary sequences, CAP generalizes to any protein-nucleic acid pair with minimal preprocessing, enabling genome-scale analyses without 3D structural data and promising faster virtual screening for drug discovery and protein engineering.
format Preprint
id arxiv_https___arxiv_org_abs_2510_22130
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CAP: Commutative Algebra Prediction of Protein-Nucleic Acid Binding Affinities
Zia, Mushal
Suwayyid, Faisal
Hozumi, Yuta
Wee, JunJie
Feng, Hongsong
Wei, Guo-Wei
Quantitative Methods
An accurate prediction of protein-nucleic acid binding affinity is vital for deciphering genomic processes, yet existing approaches often struggle in reconciling high accuracy with interpretability and computational efficiency. In this study, we introduce commutative algebra prediction (CAP), which couples persistent Stanley-Reisner theory with advanced sequence embedding for predicting protein-nucleic acid binding affinities. CAP encodes proteins through transformer-learned embeddings that retain long-range evolutionary context and represents DNA and RNA with $\textit{k}$-mer algebra embeddings derived from persistent facet ideals, which capture fine-scale nucleotide geometry. We demonstrate that CAP surpasses the SVSBI protein-nucleic acid benchmark and, in a further test, maintains reasonable performance on newly curated protein-RNA and protein-nucleic acid datasets. Leveraging only primary sequences, CAP generalizes to any protein-nucleic acid pair with minimal preprocessing, enabling genome-scale analyses without 3D structural data and promising faster virtual screening for drug discovery and protein engineering.
title CAP: Commutative Algebra Prediction of Protein-Nucleic Acid Binding Affinities
topic Quantitative Methods
url https://arxiv.org/abs/2510.22130