deMoraes-Lab/Pandoomain: Pandoomain v1.0.0 - Initial Release

Fuente: Zenodo
Enregistré dans:
Détails bibliographiques
Auteur principal: dd Emanuel
Format: Recurso digital
Publié: Zenodo 2025
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866902126819540992
author dd Emanuel
author_facet dd Emanuel
contents <p>We are excited to announce the first official release of Pandoomain, a powerful and automated Snakemake pipeline for analyzing protein domain architecture and genomic context across thousands of bacterial genomes.</p> <p>This initial release (v1.0.0) provides a robust framework for large-scale, context-driven functional genomics, enabling researchers to uncover novel protein functions, discover new toxin systems, and explore the evolutionary patterns of protein domains.</p> <p><strong>Key Features</strong></p> <ul> <li><p>Automated Data Acquisition: Pandoomain streamlines your workflow by starting with just a list of NCBI RefSeq genome assembly accessions. It automatically downloads all required protein FASTA and annotation (GFF) files, making large-scale analysis accessible.</p> </li> <li><p>HMM-Based Protein Identification: Utilizes HMMER to accurately identify target protein families using user-provided HMM profiles, allowing for highly customizable and sensitive searches.</p> </li> <li><p>Genomic Context Analysis: A core feature of the pipeline is the automated analysis of the genomic neighborhood surrounding each identified protein. This leverages the "guilt by association" principle to make powerful functional inferences.</p> </li> <li><p>Comprehensive Domain Annotation: Integrates InterProScan to perform large-scale protein domain annotation on both the initial hits and their neighbors, providing a complete picture of the domain landscape.</p> </li> <li><p>Standardized Taxonomy: Incorporates the Taxallnomy tool to generate a hierarchically complete taxonomic table, resolving a common challenge in large-scale comparative genomics and enabling robust analysis of domain distribution across the bacterial tree of life.</p> </li> <li><p>Scalable & Reproducible: Built with Snakemake, the pipeline is inherently scalable, reproducible, and fault-tolerant. It can be run locally on a powerful machine and can seamlessly restart after an interruption.</p> </li> <li><p>Structured Outputs: Generates a series of well-organized, plain-text files (TSV and FASTA) that are easy to parse, process, and use for downstream analysis and visualization.</p> </li> </ul>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_17108890
institution Zenodo
language
publishDate 2025
publisher Zenodo
record_format zenodo
spellingShingle deMoraes-Lab/Pandoomain: Pandoomain v1.0.0 - Initial Release
dd Emanuel
<p>We are excited to announce the first official release of Pandoomain, a powerful and automated Snakemake pipeline for analyzing protein domain architecture and genomic context across thousands of bacterial genomes.</p> <p>This initial release (v1.0.0) provides a robust framework for large-scale, context-driven functional genomics, enabling researchers to uncover novel protein functions, discover new toxin systems, and explore the evolutionary patterns of protein domains.</p> <p><strong>Key Features</strong></p> <ul> <li><p>Automated Data Acquisition: Pandoomain streamlines your workflow by starting with just a list of NCBI RefSeq genome assembly accessions. It automatically downloads all required protein FASTA and annotation (GFF) files, making large-scale analysis accessible.</p> </li> <li><p>HMM-Based Protein Identification: Utilizes HMMER to accurately identify target protein families using user-provided HMM profiles, allowing for highly customizable and sensitive searches.</p> </li> <li><p>Genomic Context Analysis: A core feature of the pipeline is the automated analysis of the genomic neighborhood surrounding each identified protein. This leverages the "guilt by association" principle to make powerful functional inferences.</p> </li> <li><p>Comprehensive Domain Annotation: Integrates InterProScan to perform large-scale protein domain annotation on both the initial hits and their neighbors, providing a complete picture of the domain landscape.</p> </li> <li><p>Standardized Taxonomy: Incorporates the Taxallnomy tool to generate a hierarchically complete taxonomic table, resolving a common challenge in large-scale comparative genomics and enabling robust analysis of domain distribution across the bacterial tree of life.</p> </li> <li><p>Scalable & Reproducible: Built with Snakemake, the pipeline is inherently scalable, reproducible, and fault-tolerant. It can be run locally on a powerful machine and can seamlessly restart after an interruption.</p> </li> <li><p>Structured Outputs: Generates a series of well-organized, plain-text files (TSV and FASTA) that are easy to parse, process, and use for downstream analysis and visualization.</p> </li> </ul>
title deMoraes-Lab/Pandoomain: Pandoomain v1.0.0 - Initial Release
url https://doi.org/10.5281/zenodo.17108890