DiMA: Sequence Diversity Dynamics Analyser for Viruses

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Tharanga, Shan, Unlu, Eyyub Selim, Hu, Yongli, Sjaugi, Muhammad Farhan, Celik, Muhammet A., Hekimoglu, Hilal, Miotto, Olivo, Oncel, Muhammed Miran, Khan, Asif M.
Formato: Preprint
Publicado: 2022
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913447600455680
author Tharanga, Shan
Unlu, Eyyub Selim
Hu, Yongli
Sjaugi, Muhammad Farhan
Celik, Muhammet A.
Hekimoglu, Hilal
Miotto, Olivo
Oncel, Muhammed Miran
Khan, Asif M.
author_facet Tharanga, Shan
Unlu, Eyyub Selim
Hu, Yongli
Sjaugi, Muhammad Farhan
Celik, Muhammet A.
Hekimoglu, Hilal
Miotto, Olivo
Oncel, Muhammed Miran
Khan, Asif M.
contents Sequence diversity is one of the major challenges in the design of diagnostic, prophylactic and therapeutic interventions against viruses. DiMA is a novel tool that is big data-ready and designed to facilitate the dissection of sequence diversity dynamics for viruses. DiMA stands out from other diversity analysis tools by offering various unique features. DiMA provides a quantitative overview of sequence (nucleotide/protein) diversity by use of Shannon's entropy corrected for size bias, applied via a user-defined k-mer sliding window to an input alignment file, and each k-mer position is dissected to various diversity motifs. The motifs are defined based on the probability of distinct sequences at a given k-mer position, whereby an index is the predominant sequence, while all the others are (total) variants to the index. The total variants are sub-classified into the major (most common) variant, minor variants (occurring more than once and of frequency lower than the major), and the unique (singleton) variants. DiMA allows user-defined, sequence metadata enrichment for analyses of the motifs. The application of DiMA was demonstrated for the alignment data of the relatively conserved Spike protein (2,106,985 sequences) of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) and the relatively highly diverse Pol protein (3,874) of human immunodeficiency virus-1 (HIV-1). The tool is publicly available as a web server (https://dima.bezmialem.edu.tr), as a Python library (via PyPi) and as a command line client (Via GitHub).
format Preprint
id arxiv_https___arxiv_org_abs_2205_13915
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle DiMA: Sequence Diversity Dynamics Analyser for Viruses
Tharanga, Shan
Unlu, Eyyub Selim
Hu, Yongli
Sjaugi, Muhammad Farhan
Celik, Muhammet A.
Hekimoglu, Hilal
Miotto, Olivo
Oncel, Muhammed Miran
Khan, Asif M.
Genomics
Quantitative Methods
Sequence diversity is one of the major challenges in the design of diagnostic, prophylactic and therapeutic interventions against viruses. DiMA is a novel tool that is big data-ready and designed to facilitate the dissection of sequence diversity dynamics for viruses. DiMA stands out from other diversity analysis tools by offering various unique features. DiMA provides a quantitative overview of sequence (nucleotide/protein) diversity by use of Shannon's entropy corrected for size bias, applied via a user-defined k-mer sliding window to an input alignment file, and each k-mer position is dissected to various diversity motifs. The motifs are defined based on the probability of distinct sequences at a given k-mer position, whereby an index is the predominant sequence, while all the others are (total) variants to the index. The total variants are sub-classified into the major (most common) variant, minor variants (occurring more than once and of frequency lower than the major), and the unique (singleton) variants. DiMA allows user-defined, sequence metadata enrichment for analyses of the motifs. The application of DiMA was demonstrated for the alignment data of the relatively conserved Spike protein (2,106,985 sequences) of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) and the relatively highly diverse Pol protein (3,874) of human immunodeficiency virus-1 (HIV-1). The tool is publicly available as a web server (https://dima.bezmialem.edu.tr), as a Python library (via PyPi) and as a command line client (Via GitHub).
title DiMA: Sequence Diversity Dynamics Analyser for Viruses
topic Genomics
Quantitative Methods
url https://arxiv.org/abs/2205.13915