Saved in:
| Main Author: | |
|---|---|
| Format: | Recurso digital |
| Language: | |
| Published: |
Zenodo
2025
|
| Online Access: | https://doi.org/10.5281/zenodo.18064127 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866901479605927936 |
|---|---|
| author | Laboratory of Applied Research in Mycobacteria (LaPAM) |
| author_facet | Laboratory of Applied Research in Mycobacteria (LaPAM) |
| contents | <p><strong>Kaiju database – Mycobacterium pre-compiled subset (2024 release)</strong></p> <p><strong><br></strong>This dataset provides a<strong> pre-compiled Kaiju database </strong>containing protein sequences exclusively from the genus <strong>Mycobacterium</strong>, extracted from the NCBI NR/RefSeq repositories (August 2024).</p> <p>The database was built to <strong>optimize the taxonomic classification of sequencing reads from <em>Mycobacterium tuberculosis</em> and related species</strong>, significantly reducing computational requirements compared to the full Kaiju NR database (~100 GB).<br><br>Unlike the standard Kaiju NR database or raw FASTA-based subsets, this release distributes the final Kaiju index files already built, allowing immediate use in analysis pipelines without requiring database construction.</p> <p>This subset includes representative genomes from <em>Mycobacterium tuberculosis</em>, <em>M. bovis</em>, <em>M. africanum</em>, <em>M. smegmatis</em>, and other clinically or environmentally relevant species within the genus.</p> <p><strong>Contents:</strong></p> <ul> <li> <p><code>kaiju_db_mycobacterium_2024.fmi</code> — Kaiju formatted database index</p> </li> <li> <p><code>nodes.dmp</code>, <code>names.dmp</code> — NCBI taxonomy mapping files</p> </li> </ul> <p><strong>Total size:</strong> ~1 GB<br><strong>Kaiju version:</strong> compatible with ≥ 1.9.0<br><strong>Reference source:</strong> NCBI NR/RefSeq (retrieved August 2024)</p> <p><strong>Use case:</strong><br>Designed for pipelines performing <strong>taxonomic classification and contamination screening of <em>Mycobacterium</em> sequencing data</strong>, enabling faster execution while maintaining taxonomic resolution at the species level.</p> <p><strong>Recommended citation:</strong></p> <p><em>Kaiju database – Mycobacterium subset (2024 release).</em> Zenodo. https://10.5281/zenodo.17554952</p> <p><strong>Menzel, P., Ng, K. L., & Krogh, A.</strong> (2016). <em>Fast and sensitive taxonomic classification for metagenomics with Kaiju.</em> <strong>Nature Communications</strong>, 7, 11257. <a target="_new" rel="noopener">https://doi.org/10.1038/ncomms11257</a></p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_18064127 |
| institution | Zenodo |
| language | |
| publishDate | 2025 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | kaiju_mycobacterium_pre-compiled Laboratory of Applied Research in Mycobacteria (LaPAM) <p><strong>Kaiju database – Mycobacterium pre-compiled subset (2024 release)</strong></p> <p><strong><br></strong>This dataset provides a<strong> pre-compiled Kaiju database </strong>containing protein sequences exclusively from the genus <strong>Mycobacterium</strong>, extracted from the NCBI NR/RefSeq repositories (August 2024).</p> <p>The database was built to <strong>optimize the taxonomic classification of sequencing reads from <em>Mycobacterium tuberculosis</em> and related species</strong>, significantly reducing computational requirements compared to the full Kaiju NR database (~100 GB).<br><br>Unlike the standard Kaiju NR database or raw FASTA-based subsets, this release distributes the final Kaiju index files already built, allowing immediate use in analysis pipelines without requiring database construction.</p> <p>This subset includes representative genomes from <em>Mycobacterium tuberculosis</em>, <em>M. bovis</em>, <em>M. africanum</em>, <em>M. smegmatis</em>, and other clinically or environmentally relevant species within the genus.</p> <p><strong>Contents:</strong></p> <ul> <li> <p><code>kaiju_db_mycobacterium_2024.fmi</code> — Kaiju formatted database index</p> </li> <li> <p><code>nodes.dmp</code>, <code>names.dmp</code> — NCBI taxonomy mapping files</p> </li> </ul> <p><strong>Total size:</strong> ~1 GB<br><strong>Kaiju version:</strong> compatible with ≥ 1.9.0<br><strong>Reference source:</strong> NCBI NR/RefSeq (retrieved August 2024)</p> <p><strong>Use case:</strong><br>Designed for pipelines performing <strong>taxonomic classification and contamination screening of <em>Mycobacterium</em> sequencing data</strong>, enabling faster execution while maintaining taxonomic resolution at the species level.</p> <p><strong>Recommended citation:</strong></p> <p><em>Kaiju database – Mycobacterium subset (2024 release).</em> Zenodo. https://10.5281/zenodo.17554952</p> <p><strong>Menzel, P., Ng, K. L., & Krogh, A.</strong> (2016). <em>Fast and sensitive taxonomic classification for metagenomics with Kaiju.</em> <strong>Nature Communications</strong>, 7, 11257. <a target="_new" rel="noopener">https://doi.org/10.1038/ncomms11257</a></p> |
| title | kaiju_mycobacterium_pre-compiled |
| url | https://doi.org/10.5281/zenodo.18064127 |