Material for LORA paper (Veillonella Parvula case)

Fuente: Zenodo
Salvato in:
Dettagli Bibliografici
Autore principale: Cokelaer, Thomas
Natura: Recurso digital
Pubblicazione: Zenodo 2024
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866902109519085568
author Cokelaer, Thomas
author_facet Cokelaer, Thomas
contents <h1><strong>Veillonella dataset</strong></h1> <p>LORA is a LO Read assembly pipeline part of Sequana (<a href="https://github.com/sequana/lora">https://github.com/sequana/lora</a>)</p> <p>This dataset corresponds to the <em>Veillonella</em> isolate used in the <strong>LORA</strong> paper to illustrate bacterial genome assembly from long-read sequencing data.</p> <p>The raw sequencing data originate from the study available at:<br> <a href="https://pubmed.ncbi.nlm.nih.gov/32817093/" target="_blank" rel="noopener">PubMed 32817093</a></p> <p>Specifically, the dataset is associated with the SRA run:</p> <p><a href="https://trace.ncbi.nlm.nih.gov/Traces/?view=run_browser&acc=ERR3958992&display=metadata" target="_blank" rel="noopener">ERR3958992</a></p> <p>In the original study, sequencing was performed on a <strong>PacBio Sequel II</strong> platform. The available files in the SRA project contain only <strong>raw subreads</strong> in FASTQ format. According to the authors, subreads were used without CCS construction. </p> <p>For the <strong>LORA</strong> paper, we aimed to generate <strong>higher-quality CCS reads</strong> using a <strong>minimum of 3 passes</strong> and <strong>RQ ≥ 70%</strong>, which required access to the original <strong>subreads BAM file</strong>. This BAM file was kindly provided by the authors and is shared here to ensure full reproducibility of our analysis.</p> <p>For convenience, we provide the following files:</p> <ul> <li> <p><strong><code>m54091_180306_141024.subreads.fastq.gz</code></strong> — FASTQ file equivalent to the SRA data from the publication. This file contains <strong>338,310 reads</strong>.</p> </li> <li> <p><strong><code>veillonella.subreads.bam</code></strong> — Original PacBio subreads in BAM format, which generated the data referenced in the  publication.</p> </li> <li> <p><strong><code>veillonella.ccs.bam</code></strong> — CCS (Circular Consensus Sequencing) reads generated from the subreads with the following parameters:</p> <ul> <li> <p><strong>Minimum read quality (RQ): 0.7</strong></p> </li> <li> <p><strong>Minimum number of passes: 0</strong><br>These settings produce CCS reads with relatively few passes.</p> </li> </ul> <p> </p> </li> <li> <p><strong><code>veillonella.ccs.fastq.gz</code></strong> — FASTQ version of the above CCS dataset, provided for convenience.</p> </li> </ul> <p>These data files enable reproducibility of the analyses presented in the LORA paper and can be reused for independent benchmarking or assembly studies.</p> <p> </p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_13306684
institution Zenodo
language
publishDate 2024
publisher Zenodo
record_format zenodo
spellingShingle Material for LORA paper (Veillonella Parvula case)
Cokelaer, Thomas
<h1><strong>Veillonella dataset</strong></h1> <p>LORA is a LO Read assembly pipeline part of Sequana (<a href="https://github.com/sequana/lora">https://github.com/sequana/lora</a>)</p> <p>This dataset corresponds to the <em>Veillonella</em> isolate used in the <strong>LORA</strong> paper to illustrate bacterial genome assembly from long-read sequencing data.</p> <p>The raw sequencing data originate from the study available at:<br> <a href="https://pubmed.ncbi.nlm.nih.gov/32817093/" target="_blank" rel="noopener">PubMed 32817093</a></p> <p>Specifically, the dataset is associated with the SRA run:</p> <p><a href="https://trace.ncbi.nlm.nih.gov/Traces/?view=run_browser&acc=ERR3958992&display=metadata" target="_blank" rel="noopener">ERR3958992</a></p> <p>In the original study, sequencing was performed on a <strong>PacBio Sequel II</strong> platform. The available files in the SRA project contain only <strong>raw subreads</strong> in FASTQ format. According to the authors, subreads were used without CCS construction. </p> <p>For the <strong>LORA</strong> paper, we aimed to generate <strong>higher-quality CCS reads</strong> using a <strong>minimum of 3 passes</strong> and <strong>RQ ≥ 70%</strong>, which required access to the original <strong>subreads BAM file</strong>. This BAM file was kindly provided by the authors and is shared here to ensure full reproducibility of our analysis.</p> <p>For convenience, we provide the following files:</p> <ul> <li> <p><strong><code>m54091_180306_141024.subreads.fastq.gz</code></strong> — FASTQ file equivalent to the SRA data from the publication. This file contains <strong>338,310 reads</strong>.</p> </li> <li> <p><strong><code>veillonella.subreads.bam</code></strong> — Original PacBio subreads in BAM format, which generated the data referenced in the  publication.</p> </li> <li> <p><strong><code>veillonella.ccs.bam</code></strong> — CCS (Circular Consensus Sequencing) reads generated from the subreads with the following parameters:</p> <ul> <li> <p><strong>Minimum read quality (RQ): 0.7</strong></p> </li> <li> <p><strong>Minimum number of passes: 0</strong><br>These settings produce CCS reads with relatively few passes.</p> </li> </ul> <p> </p> </li> <li> <p><strong><code>veillonella.ccs.fastq.gz</code></strong> — FASTQ version of the above CCS dataset, provided for convenience.</p> </li> </ul> <p>These data files enable reproducibility of the analyses presented in the LORA paper and can be reused for independent benchmarking or assembly studies.</p> <p> </p>
title Material for LORA paper (Veillonella Parvula case)
url https://doi.org/10.5281/zenodo.13306684