Calculating complexity of large randomized libraries

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autore principale: Kong, Yong
Natura: Preprint
Pubblicazione: 2014
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914810445168640
author Kong, Yong
author_facet Kong, Yong
contents Randomized libraries are increasingly popular in protein engineering and other biomedical research fields. Statistics of the libraries are useful to guide and evaluate randomized library construction. Previous works only give the mean of the number of unique sequences in the library, and they can only handle equal molar ratio of the four nucleotides at a small number of mutation sites. We derive formulas to calculate the mean and variance of the number of unique sequences in libraries generated by cassette mutagenesis with mixtures of arbitrary nucleotide ratios. Computer program was developed which utilizes arbitrary numerical precision software package to calculate the statistics of large libraries. The statistics of library with mutations in more than $20$ amino acids can be calculated easily. Results show that the nucleotide ratios have significant effects on these statistics. The more skewed the ratio, the larger the library size is needed to obtain the same expected number of unique sequences. The program is freely available at \url{http://graphics.med.yale.edu/cgi-bin/lib_comp.pl}.
format Preprint
id arxiv_https___arxiv_org_abs_1410_5851
institution arXiv
publishDate 2014
record_format arxiv
spellingShingle Calculating complexity of large randomized libraries
Kong, Yong
Quantitative Methods
Computation
Randomized libraries are increasingly popular in protein engineering and other biomedical research fields. Statistics of the libraries are useful to guide and evaluate randomized library construction. Previous works only give the mean of the number of unique sequences in the library, and they can only handle equal molar ratio of the four nucleotides at a small number of mutation sites. We derive formulas to calculate the mean and variance of the number of unique sequences in libraries generated by cassette mutagenesis with mixtures of arbitrary nucleotide ratios. Computer program was developed which utilizes arbitrary numerical precision software package to calculate the statistics of large libraries. The statistics of library with mutations in more than $20$ amino acids can be calculated easily. Results show that the nucleotide ratios have significant effects on these statistics. The more skewed the ratio, the larger the library size is needed to obtain the same expected number of unique sequences. The program is freely available at \url{http://graphics.med.yale.edu/cgi-bin/lib_comp.pl}.
title Calculating complexity of large randomized libraries
topic Quantitative Methods
Computation
url https://arxiv.org/abs/1410.5851