Saved in:
Bibliographic Details
Main Authors: Feller, Aaron L., Wilke, Claus O.
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2505.00719
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908345638584320
author Feller, Aaron L.
Wilke, Claus O.
author_facet Feller, Aaron L.
Wilke, Claus O.
contents Converting peptide sequences into useful representations for downstream analysis is a common step in computational modeling and cheminformatics. Furthermore, peptide drugs (e.g., Semaglutide, Degarelix) often take advantage of the diverse chemistries found in noncanonical amino acids (NCAAs), altered stereochemistry, and backbone modifications. Despite there being several chemoinformatics toolkits, none are tailored to the task of converting a modified peptide from an amino acid representation to the chemical string nomenclature Simplified Molecular-Input Line-Entry System (SMILES), often used in chemical modeling. Here we present p2smi, a Python toolkit with CLI, designed to facilitate the conversion of peptide sequences into chemical SMILES strings. By supporting both cyclic and linear peptides, including those with NCAAs, p2smi enables researchers to generate accurate SMILES strings for drug-like peptides, reducing the overhead for computational modeling and cheminformatics analyses. The toolkit also offers functionalities for chemical modification, synthesis feasibility evaluation, and calculation of molecular properties such as hydrophobicity, topological polar surface area, molecular weight, and adherence to Lipinski's rules for drug-likeness.
format Preprint
id arxiv_https___arxiv_org_abs_2505_00719
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle p2smi: A Python Toolkit for Peptide FASTA-to-SMILES Conversion and Molecular Property Analysis
Feller, Aaron L.
Wilke, Claus O.
Biomolecules
Converting peptide sequences into useful representations for downstream analysis is a common step in computational modeling and cheminformatics. Furthermore, peptide drugs (e.g., Semaglutide, Degarelix) often take advantage of the diverse chemistries found in noncanonical amino acids (NCAAs), altered stereochemistry, and backbone modifications. Despite there being several chemoinformatics toolkits, none are tailored to the task of converting a modified peptide from an amino acid representation to the chemical string nomenclature Simplified Molecular-Input Line-Entry System (SMILES), often used in chemical modeling. Here we present p2smi, a Python toolkit with CLI, designed to facilitate the conversion of peptide sequences into chemical SMILES strings. By supporting both cyclic and linear peptides, including those with NCAAs, p2smi enables researchers to generate accurate SMILES strings for drug-like peptides, reducing the overhead for computational modeling and cheminformatics analyses. The toolkit also offers functionalities for chemical modification, synthesis feasibility evaluation, and calculation of molecular properties such as hydrophobicity, topological polar surface area, molecular weight, and adherence to Lipinski's rules for drug-likeness.
title p2smi: A Python Toolkit for Peptide FASTA-to-SMILES Conversion and Molecular Property Analysis
topic Biomolecules
url https://arxiv.org/abs/2505.00719