EBIC: an open source software for high-dimensional and big data biclustering analyses

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Orzechowski, Patryk, Moore, Jason H.
Format: Preprint
Published: 2018
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917767611940864
author Orzechowski, Patryk
Moore, Jason H.
author_facet Orzechowski, Patryk
Moore, Jason H.
contents Motivation: In this paper we present the latest release of EBIC, a next-generation biclustering algorithm for mining genetic data. The major contribution of this paper is adding support for big data, making it possible to efficiently run large genomic data mining analyses. Additional enhancements include integration with R and Bioconductor and an option to remove influence of missing value on the final result. Results: EBIC was applied to datasets of different sizes, including a large DNA methylation dataset with 436,444 rows. For the largest dataset we observed over 6.6 fold speedup in computation time on a cluster of 8 GPUs compared to running the method on a single GPU. This proves high scalability of the algorithm. Availability: The latest version of EBIC could be downloaded from http://github.com/EpistasisLab/ebic . Installation and usage instructions are also available online.
format Preprint
id arxiv_https___arxiv_org_abs_1807_09932
institution arXiv
publishDate 2018
record_format arxiv
spellingShingle EBIC: an open source software for high-dimensional and big data biclustering analyses
Orzechowski, Patryk
Moore, Jason H.
Genomics
Machine Learning
68, 92
I.5.2; I.2.11; I.5.3; J.3; I.2.0
Motivation: In this paper we present the latest release of EBIC, a next-generation biclustering algorithm for mining genetic data. The major contribution of this paper is adding support for big data, making it possible to efficiently run large genomic data mining analyses. Additional enhancements include integration with R and Bioconductor and an option to remove influence of missing value on the final result. Results: EBIC was applied to datasets of different sizes, including a large DNA methylation dataset with 436,444 rows. For the largest dataset we observed over 6.6 fold speedup in computation time on a cluster of 8 GPUs compared to running the method on a single GPU. This proves high scalability of the algorithm. Availability: The latest version of EBIC could be downloaded from http://github.com/EpistasisLab/ebic . Installation and usage instructions are also available online.
title EBIC: an open source software for high-dimensional and big data biclustering analyses
topic Genomics
Machine Learning
68, 92
I.5.2; I.2.11; I.5.3; J.3; I.2.0
url https://arxiv.org/abs/1807.09932