Material for LORA paper (Veillonella Parvula case)
Fuente:
Zenodo
Salvato in:
| Autore principale: | |
|---|---|
| Natura: | Recurso digital |
| Pubblicazione: |
Zenodo
2024
|
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866902109519085568 |
|---|---|
| author | Cokelaer, Thomas |
| author_facet | Cokelaer, Thomas |
| contents | <h1><strong>Veillonella dataset</strong></h1> <p>LORA is a LO Read assembly pipeline part of Sequana (<a href="https://github.com/sequana/lora">https://github.com/sequana/lora</a>)</p> <p>This dataset corresponds to the <em>Veillonella</em> isolate used in the <strong>LORA</strong> paper to illustrate bacterial genome assembly from long-read sequencing data.</p> <p>The raw sequencing data originate from the study available at:<br> <a href="https://pubmed.ncbi.nlm.nih.gov/32817093/" target="_blank" rel="noopener">PubMed 32817093</a></p> <p>Specifically, the dataset is associated with the SRA run:</p> <p><a href="https://trace.ncbi.nlm.nih.gov/Traces/?view=run_browser&acc=ERR3958992&display=metadata" target="_blank" rel="noopener">ERR3958992</a></p> <p>In the original study, sequencing was performed on a <strong>PacBio Sequel II</strong> platform. The available files in the SRA project contain only <strong>raw subreads</strong> in FASTQ format. According to the authors, subreads were used without CCS construction. </p> <p>For the <strong>LORA</strong> paper, we aimed to generate <strong>higher-quality CCS reads</strong> using a <strong>minimum of 3 passes</strong> and <strong>RQ ≥ 70%</strong>, which required access to the original <strong>subreads BAM file</strong>. This BAM file was kindly provided by the authors and is shared here to ensure full reproducibility of our analysis.</p> <p>For convenience, we provide the following files:</p> <ul> <li> <p><strong><code>m54091_180306_141024.subreads.fastq.gz</code></strong> — FASTQ file equivalent to the SRA data from the publication. This file contains <strong>338,310 reads</strong>.</p> </li> <li> <p><strong><code>veillonella.subreads.bam</code></strong> — Original PacBio subreads in BAM format, which generated the data referenced in the publication.</p> </li> <li> <p><strong><code>veillonella.ccs.bam</code></strong> — CCS (Circular Consensus Sequencing) reads generated from the subreads with the following parameters:</p> <ul> <li> <p><strong>Minimum read quality (RQ): 0.7</strong></p> </li> <li> <p><strong>Minimum number of passes: 0</strong><br>These settings produce CCS reads with relatively few passes.</p> </li> </ul> <p> </p> </li> <li> <p><strong><code>veillonella.ccs.fastq.gz</code></strong> — FASTQ version of the above CCS dataset, provided for convenience.</p> </li> </ul> <p>These data files enable reproducibility of the analyses presented in the LORA paper and can be reused for independent benchmarking or assembly studies.</p> <p> </p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_13306684 |
| institution | Zenodo |
| language | |
| publishDate | 2024 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Material for LORA paper (Veillonella Parvula case) Cokelaer, Thomas <h1><strong>Veillonella dataset</strong></h1> <p>LORA is a LO Read assembly pipeline part of Sequana (<a href="https://github.com/sequana/lora">https://github.com/sequana/lora</a>)</p> <p>This dataset corresponds to the <em>Veillonella</em> isolate used in the <strong>LORA</strong> paper to illustrate bacterial genome assembly from long-read sequencing data.</p> <p>The raw sequencing data originate from the study available at:<br> <a href="https://pubmed.ncbi.nlm.nih.gov/32817093/" target="_blank" rel="noopener">PubMed 32817093</a></p> <p>Specifically, the dataset is associated with the SRA run:</p> <p><a href="https://trace.ncbi.nlm.nih.gov/Traces/?view=run_browser&acc=ERR3958992&display=metadata" target="_blank" rel="noopener">ERR3958992</a></p> <p>In the original study, sequencing was performed on a <strong>PacBio Sequel II</strong> platform. The available files in the SRA project contain only <strong>raw subreads</strong> in FASTQ format. According to the authors, subreads were used without CCS construction. </p> <p>For the <strong>LORA</strong> paper, we aimed to generate <strong>higher-quality CCS reads</strong> using a <strong>minimum of 3 passes</strong> and <strong>RQ ≥ 70%</strong>, which required access to the original <strong>subreads BAM file</strong>. This BAM file was kindly provided by the authors and is shared here to ensure full reproducibility of our analysis.</p> <p>For convenience, we provide the following files:</p> <ul> <li> <p><strong><code>m54091_180306_141024.subreads.fastq.gz</code></strong> — FASTQ file equivalent to the SRA data from the publication. This file contains <strong>338,310 reads</strong>.</p> </li> <li> <p><strong><code>veillonella.subreads.bam</code></strong> — Original PacBio subreads in BAM format, which generated the data referenced in the publication.</p> </li> <li> <p><strong><code>veillonella.ccs.bam</code></strong> — CCS (Circular Consensus Sequencing) reads generated from the subreads with the following parameters:</p> <ul> <li> <p><strong>Minimum read quality (RQ): 0.7</strong></p> </li> <li> <p><strong>Minimum number of passes: 0</strong><br>These settings produce CCS reads with relatively few passes.</p> </li> </ul> <p> </p> </li> <li> <p><strong><code>veillonella.ccs.fastq.gz</code></strong> — FASTQ version of the above CCS dataset, provided for convenience.</p> </li> </ul> <p>These data files enable reproducibility of the analyses presented in the LORA paper and can be reused for independent benchmarking or assembly studies.</p> <p> </p> |
| title | Material for LORA paper (Veillonella Parvula case) |
| url | https://doi.org/10.5281/zenodo.13306684 |