Neighborhood Stability in Double/Debiased Machine Learning with Dependent Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cao, Jianfei, Leung, Michael P.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914158328414208
author Cao, Jianfei
Leung, Michael P.
author_facet Cao, Jianfei
Leung, Michael P.
contents This paper studies double/debiased machine learning (DML) methods applied to weakly dependent data. We allow observations to be situated in a general metric space that accommodates spatial and network data. Existing work implements cross-fitting by excluding from the training fold observations sufficiently close to the evaluation fold. We find in simulations that this can result in exceedingly small training fold sizes, particularly with network data. We therefore seek to establish the validity of DML without cross-fitting, building on recent work by Chen et al. (2022). They study i.i.d. data and require the machine learner to satisfy a natural stability condition requiring insensitivity to data perturbations that resample a single observation. We extend these results to dependent data by strengthening stability to "neighborhood stability," which requires insensitivity to resampling observations in any slowly growing neighborhood. We show that existing results on the stability of various machine learners can be adapted to verify neighborhood stability.
format Preprint
id arxiv_https___arxiv_org_abs_2511_10995
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Neighborhood Stability in Double/Debiased Machine Learning with Dependent Data
Cao, Jianfei
Leung, Michael P.
Econometrics
Methodology
This paper studies double/debiased machine learning (DML) methods applied to weakly dependent data. We allow observations to be situated in a general metric space that accommodates spatial and network data. Existing work implements cross-fitting by excluding from the training fold observations sufficiently close to the evaluation fold. We find in simulations that this can result in exceedingly small training fold sizes, particularly with network data. We therefore seek to establish the validity of DML without cross-fitting, building on recent work by Chen et al. (2022). They study i.i.d. data and require the machine learner to satisfy a natural stability condition requiring insensitivity to data perturbations that resample a single observation. We extend these results to dependent data by strengthening stability to "neighborhood stability," which requires insensitivity to resampling observations in any slowly growing neighborhood. We show that existing results on the stability of various machine learners can be adapted to verify neighborhood stability.
title Neighborhood Stability in Double/Debiased Machine Learning with Dependent Data
topic Econometrics
Methodology
url https://arxiv.org/abs/2511.10995