Statistical Hypothesis Testing for Information Value (IV)

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rojas, Helder, Alvarez, Cirilo, Rojas, Nilton
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915757912227840
author Rojas, Helder
Alvarez, Cirilo
Rojas, Nilton
author_facet Rojas, Helder
Alvarez, Cirilo
Rojas, Nilton
contents Information Value (IV) is a widely used technique for feature selection prior to the modeling phase, particularly in credit scoring and related domains. However, conventional IV-based practices rely on fixed empirical thresholds, which lack statistical justification and may be sensitive to characteristics such as class imbalance. In this work, we develop a formal statistical framework for IV by establishing its connection with Jeffreys divergence and propose a novel nonparametric hypothesis test, referred to as the J-Divergence test. Our method provides rigorous asymptotic guarantees and enables interpretable decisions based on \(p\)-values. Numerical experiments, including synthetic and real-world data, demonstrate that the proposed test is more reliable than traditional IV thresholding, particularly under strong imbalance. The test is model-agnostic, computationally efficient, and well-suited for the pre-modeling phase in high-dimensional or imbalanced settings. An open-source Python library is provided for reproducibility and practical adoption.
format Preprint
id arxiv_https___arxiv_org_abs_2309_13183
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Statistical Hypothesis Testing for Information Value (IV)
Rojas, Helder
Alvarez, Cirilo
Rojas, Nilton
Statistics Theory
Methodology
Machine Learning
Information Value (IV) is a widely used technique for feature selection prior to the modeling phase, particularly in credit scoring and related domains. However, conventional IV-based practices rely on fixed empirical thresholds, which lack statistical justification and may be sensitive to characteristics such as class imbalance. In this work, we develop a formal statistical framework for IV by establishing its connection with Jeffreys divergence and propose a novel nonparametric hypothesis test, referred to as the J-Divergence test. Our method provides rigorous asymptotic guarantees and enables interpretable decisions based on \(p\)-values. Numerical experiments, including synthetic and real-world data, demonstrate that the proposed test is more reliable than traditional IV thresholding, particularly under strong imbalance. The test is model-agnostic, computationally efficient, and well-suited for the pre-modeling phase in high-dimensional or imbalanced settings. An open-source Python library is provided for reproducibility and practical adoption.
title Statistical Hypothesis Testing for Information Value (IV)
topic Statistics Theory
Methodology
Machine Learning
url https://arxiv.org/abs/2309.13183