Chemical Reaction Optimization for Protein Complex Prediction in Large Protein-Protein Interaction Network

Fuente: Zenodo
Gespeichert in:
Bibliographische Detailangaben
1. Verfasser: Islam, Md. Shahidul
Format: Recurso digital
Sprache:Englisch
Veröffentlicht: Zenodo 2021
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866901248438960128
author Islam, Md. Shahidul
author_facet Islam, Md. Shahidul
contents <div> <p>Due to high computational complexity, the detection of protein complexes in large protein-protein interaction (PPI) networks remains a challenging problem. Finding the actual protein complexes from a large PPI network requires a sophisticated algorithm to handle the complexity. The protein complexes are exhibited in densely connected sub-graphs in a PPI network. This paper presents a novel algorithm based on a metaheuristic method for protein complex prediction in large PPI networks. The algorithm mimics the density-based graph clustering method with biological heuristics to identify the protein complexes. A PPI network can be transcribed as an undirected graph where a node represents a protein and an edge represents an interaction between two proteins. The algorithm works in three main steps: identifying the core proteins (seed), propagating core complexes, and optimizing the complexes. A local walk algorithm called probabilistic local walks (PLW) is used to score and extract seeds. A vertex scoring function defined as the product of the degree of a vertex and the density of its sub-graph is used during seed selection. The top 30% of vertices are considered core proteins. Such a scoring function has two significant advantages. Between two proteins with the same sub-graph density yet different degrees, the protein with a higher degree is more likely to be selected as a core protein. On the other side, two proteins with the same degree but different density, the protein with higher sub-graph density tends to be selected as the seed protein. Therefore, a balanced measure between degree and density helps to select the optimal seed proteins. The algorithm visits neighboring nodes of each seed node. The frequency of visiting nodes from a seed node is counted, and a z-score is computed for non-zero frequencies. The proteins above a significance level are chosen as complex cores. CRO takes over the complex cores and optimizes them to the final protein complexes. The operators of the CRO algorithm are redesigned, and the parameters are tuned to find a suitable combination for the protein complex prediction problem. Additionally, three more repair operators are designed to improve the performance of the CRO algorithm. The method was applied to the yeast protein interaction data and compared with the state-of-the-art algorithms. The comparisons demonstrate the best performance of the proposed algorithm in terms of Accuracy and F-measure.</p> </div> <div> <div> </div> </div>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_17164020
institution Zenodo
language eng
publishDate 2021
publisher Zenodo
record_format zenodo
spellingShingle Chemical Reaction Optimization for Protein Complex Prediction in Large Protein-Protein Interaction Network
Islam, Md. Shahidul
Bioinformatics
Protein Clustering
<div> <p>Due to high computational complexity, the detection of protein complexes in large protein-protein interaction (PPI) networks remains a challenging problem. Finding the actual protein complexes from a large PPI network requires a sophisticated algorithm to handle the complexity. The protein complexes are exhibited in densely connected sub-graphs in a PPI network. This paper presents a novel algorithm based on a metaheuristic method for protein complex prediction in large PPI networks. The algorithm mimics the density-based graph clustering method with biological heuristics to identify the protein complexes. A PPI network can be transcribed as an undirected graph where a node represents a protein and an edge represents an interaction between two proteins. The algorithm works in three main steps: identifying the core proteins (seed), propagating core complexes, and optimizing the complexes. A local walk algorithm called probabilistic local walks (PLW) is used to score and extract seeds. A vertex scoring function defined as the product of the degree of a vertex and the density of its sub-graph is used during seed selection. The top 30% of vertices are considered core proteins. Such a scoring function has two significant advantages. Between two proteins with the same sub-graph density yet different degrees, the protein with a higher degree is more likely to be selected as a core protein. On the other side, two proteins with the same degree but different density, the protein with higher sub-graph density tends to be selected as the seed protein. Therefore, a balanced measure between degree and density helps to select the optimal seed proteins. The algorithm visits neighboring nodes of each seed node. The frequency of visiting nodes from a seed node is counted, and a z-score is computed for non-zero frequencies. The proteins above a significance level are chosen as complex cores. CRO takes over the complex cores and optimizes them to the final protein complexes. The operators of the CRO algorithm are redesigned, and the parameters are tuned to find a suitable combination for the protein complex prediction problem. Additionally, three more repair operators are designed to improve the performance of the CRO algorithm. The method was applied to the yeast protein interaction data and compared with the state-of-the-art algorithms. The comparisons demonstrate the best performance of the proposed algorithm in terms of Accuracy and F-measure.</p> </div> <div> <div> </div> </div>
title Chemical Reaction Optimization for Protein Complex Prediction in Large Protein-Protein Interaction Network
topic Bioinformatics
Protein Clustering
url https://doi.org/10.5281/zenodo.17164020