Optimizing STAR Aligner for High Throughput Computing in the Cloud
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914944857931776 |
|---|---|
| author | Kica, Piotr Lichołai, Sabina Orzechowski, Michał Malawski, Maciej |
| author_facet | Kica, Piotr Lichołai, Sabina Orzechowski, Michał Malawski, Maciej |
| contents | We propose a scalable, cloud-native architecture designed for Transcriptomics Atlas Pipeline, using a resource-intensive STAR aligner and processing tens or hundreds of terabytes of RNA-seq data. We implement the pipeline using AWS cloud services, introduce performance optimizations and perform experimental evaluation in the cloud. Our optimization techniques result in computational savings thanks to the "early stopping" approach, selection of right-sized resources, and using newer version of Ensembl genome. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2409_05886 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Optimizing STAR Aligner for High Throughput Computing in the Cloud Kica, Piotr Lichołai, Sabina Orzechowski, Michał Malawski, Maciej Distributed, Parallel, and Cluster Computing We propose a scalable, cloud-native architecture designed for Transcriptomics Atlas Pipeline, using a resource-intensive STAR aligner and processing tens or hundreds of terabytes of RNA-seq data. We implement the pipeline using AWS cloud services, introduce performance optimizations and perform experimental evaluation in the cloud. Our optimization techniques result in computational savings thanks to the "early stopping" approach, selection of right-sized resources, and using newer version of Ensembl genome. |
| title | Optimizing STAR Aligner for High Throughput Computing in the Cloud |
| topic | Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2409.05886 |