Fast randomized algorithms for low-rank matrix approximations with applications in global comparative analysis of a class of data sets

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xu, Weiwei, Shen, Weijie, Li, Wen, Gao, Weiguo, Li, Yingzhou
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916205643694080
author Xu, Weiwei
Shen, Weijie
Li, Wen
Gao, Weiguo
Li, Yingzhou
author_facet Xu, Weiwei
Shen, Weijie
Li, Wen
Gao, Weiguo
Li, Yingzhou
contents Generalized singular values (GSVs) play an essential role in the comparative analysis. In the real world data for comparative analysis, both data matrices are usually numerically low-rank. This paper proposes a randomized algorithm to first approximately extract bases and then calculate GSVs efficiently. The accuracy of both basis extration and comparative analysis quantities, angular distances, generalized fractions of the eigenexpression, and generalized normalized Shannon entropy, are rigursly analyzed. The proposed algorithm is applied to both synthetic data sets and the genome-scale expression data sets. Comparing to other GSVs algorithms, the proposed algorithm achieves the fastest runtime while preserving sufficient accuracy in comparative analysis.
format Preprint
id arxiv_https___arxiv_org_abs_2404_09459
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Fast randomized algorithms for low-rank matrix approximations with applications in global comparative analysis of a class of data sets
Xu, Weiwei
Shen, Weijie
Li, Wen
Gao, Weiguo
Li, Yingzhou
Numerical Analysis
Generalized singular values (GSVs) play an essential role in the comparative analysis. In the real world data for comparative analysis, both data matrices are usually numerically low-rank. This paper proposes a randomized algorithm to first approximately extract bases and then calculate GSVs efficiently. The accuracy of both basis extration and comparative analysis quantities, angular distances, generalized fractions of the eigenexpression, and generalized normalized Shannon entropy, are rigursly analyzed. The proposed algorithm is applied to both synthetic data sets and the genome-scale expression data sets. Comparing to other GSVs algorithms, the proposed algorithm achieves the fastest runtime while preserving sufficient accuracy in comparative analysis.
title Fast randomized algorithms for low-rank matrix approximations with applications in global comparative analysis of a class of data sets
topic Numerical Analysis
url https://arxiv.org/abs/2404.09459