A Compounded Burr Probability Distribution for Fitting Heavy-Tailed Data with Applications to Biological Networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chakraborty, Tanujit, Chattopadhyay, Swarup, Das, Suchismita, Naik, Shraddha M., Hens, Chittaranjan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909729770438656
author Chakraborty, Tanujit
Chattopadhyay, Swarup
Das, Suchismita
Naik, Shraddha M.
Hens, Chittaranjan
author_facet Chakraborty, Tanujit
Chattopadhyay, Swarup
Das, Suchismita
Naik, Shraddha M.
Hens, Chittaranjan
contents Complex biological networks, encompassing metabolic pathways, gene regulatory systems, and protein-protein interaction networks, often exhibit scale-free structures characterized by heavy-tailed degree distributions. However, empirical studies reveal significant deviations from ideal power law behavior, underscoring the need for more flexible and accurate probabilistic models. In this work, we propose the Compounded Burr (CBurr) distribution, a novel four parameter family derived by compounding the Burr distribution with a discrete mixing process. This model is specifically designed to capture both the body and tail behavior of real-world network degree distributions with applications to biological networks. We rigorously derive its statistical properties, including moments, hazard and risk functions, and tail behavior, and develop an efficient maximum likelihood estimation framework. The CBurr model demonstrates broad applicability to networks with complex connectivity patterns, particularly in biological, social, and technological domains. Extensive experiments on large-scale biological network datasets show that CBurr consistently outperforms classical power-law, log-normal, and other heavy-tailed models across the full degree spectrum. By providing a statistically grounded and interpretable framework, the CBurr model enhances our ability to characterize the structural heterogeneity of biological networks.
format Preprint
id arxiv_https___arxiv_org_abs_2407_04465
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Compounded Burr Probability Distribution for Fitting Heavy-Tailed Data with Applications to Biological Networks
Chakraborty, Tanujit
Chattopadhyay, Swarup
Das, Suchismita
Naik, Shraddha M.
Hens, Chittaranjan
Applications
Social and Information Networks
Data Analysis, Statistics and Probability
Complex biological networks, encompassing metabolic pathways, gene regulatory systems, and protein-protein interaction networks, often exhibit scale-free structures characterized by heavy-tailed degree distributions. However, empirical studies reveal significant deviations from ideal power law behavior, underscoring the need for more flexible and accurate probabilistic models. In this work, we propose the Compounded Burr (CBurr) distribution, a novel four parameter family derived by compounding the Burr distribution with a discrete mixing process. This model is specifically designed to capture both the body and tail behavior of real-world network degree distributions with applications to biological networks. We rigorously derive its statistical properties, including moments, hazard and risk functions, and tail behavior, and develop an efficient maximum likelihood estimation framework. The CBurr model demonstrates broad applicability to networks with complex connectivity patterns, particularly in biological, social, and technological domains. Extensive experiments on large-scale biological network datasets show that CBurr consistently outperforms classical power-law, log-normal, and other heavy-tailed models across the full degree spectrum. By providing a statistically grounded and interpretable framework, the CBurr model enhances our ability to characterize the structural heterogeneity of biological networks.
title A Compounded Burr Probability Distribution for Fitting Heavy-Tailed Data with Applications to Biological Networks
topic Applications
Social and Information Networks
Data Analysis, Statistics and Probability
url https://arxiv.org/abs/2407.04465