Beyond Normality: Reliable A/B Testing with Non-Gaussian Data

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gong, Junpeng, Wang, Chunkai, Li, Hao, Ma, Jinyong, Li, Haoxuan, He, Xu
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918173526196224
author Gong, Junpeng
Wang, Chunkai
Li, Hao
Ma, Jinyong
Li, Haoxuan
He, Xu
author_facet Gong, Junpeng
Wang, Chunkai
Li, Hao
Ma, Jinyong
Li, Haoxuan
He, Xu
contents A/B testing has become the cornerstone of decision-making in online markets, guiding how platforms launch new features, optimize pricing strategies, and improve user experience. In practice, we typically employ the pairwise $t$-test to compare outcomes between the treatment and control groups, thereby assessing the effectiveness of a given strategy. To be trustworthy, these experiments must keep Type I error (i.e., false positive rate) under control; otherwise, we may launch harmful strategies. However, in real-world applications, we find that A/B testing often fails to deliver reliable results. When the data distribution departs from normality or when the treatment and control groups differ in sample size, the commonly used pairwise $t$-test is no longer trustworthy. In this paper, we quantify how skewed, long tailed data and unequal allocation distort error rates and derive explicit formulas for the minimum sample size required for the $t$-test to remain valid. We find that many online feedback metrics require hundreds of millions samples to ensure reliable A/B testing. Thus we introduce an Edgeworth-based correction that provides more accurate $p$-values when the available sample size is limited. Offline experiments on a leading A/B testing platform corroborate the practical value of our theoretical minimum sample size thresholds and demonstrate that the corrected method substantially improves the reliability of A/B testing in real-world conditions.
format Preprint
id arxiv_https___arxiv_org_abs_2510_23666
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Normality: Reliable A/B Testing with Non-Gaussian Data
Gong, Junpeng
Wang, Chunkai
Li, Hao
Ma, Jinyong
Li, Haoxuan
He, Xu
Machine Learning
Methodology
I.2.6; G.3; I.5.1
A/B testing has become the cornerstone of decision-making in online markets, guiding how platforms launch new features, optimize pricing strategies, and improve user experience. In practice, we typically employ the pairwise $t$-test to compare outcomes between the treatment and control groups, thereby assessing the effectiveness of a given strategy. To be trustworthy, these experiments must keep Type I error (i.e., false positive rate) under control; otherwise, we may launch harmful strategies. However, in real-world applications, we find that A/B testing often fails to deliver reliable results. When the data distribution departs from normality or when the treatment and control groups differ in sample size, the commonly used pairwise $t$-test is no longer trustworthy. In this paper, we quantify how skewed, long tailed data and unequal allocation distort error rates and derive explicit formulas for the minimum sample size required for the $t$-test to remain valid. We find that many online feedback metrics require hundreds of millions samples to ensure reliable A/B testing. Thus we introduce an Edgeworth-based correction that provides more accurate $p$-values when the available sample size is limited. Offline experiments on a leading A/B testing platform corroborate the practical value of our theoretical minimum sample size thresholds and demonstrate that the corrected method substantially improves the reliability of A/B testing in real-world conditions.
title Beyond Normality: Reliable A/B Testing with Non-Gaussian Data
topic Machine Learning
Methodology
I.2.6; G.3; I.5.1
url https://arxiv.org/abs/2510.23666