Tight Bounds for Answering Adaptively Chosen Concentrated Queries

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rapoport, Emma, Cohen, Edith, Stemmer, Uri
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914152282324992
author Rapoport, Emma
Cohen, Edith
Stemmer, Uri
author_facet Rapoport, Emma
Cohen, Edith
Stemmer, Uri
contents Most work on adaptive data analysis assumes that samples in the dataset are independent. When correlations are allowed, even the non-adaptive setting can become intractable, unless some structural constraints are imposed. To address this, Bassily and Freund [2016] introduced the elegant framework of concentrated queries, which requires the analyst to restrict itself to queries that are concentrated around their expected value. While this assumption makes the problem trivial in the non-adaptive setting, in the adaptive setting it remains quite challenging. In fact, all known algorithms in this framework support significantly fewer queries than in the independent case: At most $O(n)$ queries for a sample of size $n$, compared to $O(n^2)$ in the independent setting. In this work, we prove that this utility gap is inherent under the current formulation of the concentrated queries framework, assuming some natural conditions on the algorithm. Additionally, we present a simplified version of the best-known algorithms that match our impossibility result.
format Preprint
id arxiv_https___arxiv_org_abs_2507_13700
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Tight Bounds for Answering Adaptively Chosen Concentrated Queries
Rapoport, Emma
Cohen, Edith
Stemmer, Uri
Data Structures and Algorithms
Machine Learning
Most work on adaptive data analysis assumes that samples in the dataset are independent. When correlations are allowed, even the non-adaptive setting can become intractable, unless some structural constraints are imposed. To address this, Bassily and Freund [2016] introduced the elegant framework of concentrated queries, which requires the analyst to restrict itself to queries that are concentrated around their expected value. While this assumption makes the problem trivial in the non-adaptive setting, in the adaptive setting it remains quite challenging. In fact, all known algorithms in this framework support significantly fewer queries than in the independent case: At most $O(n)$ queries for a sample of size $n$, compared to $O(n^2)$ in the independent setting. In this work, we prove that this utility gap is inherent under the current formulation of the concentrated queries framework, assuming some natural conditions on the algorithm. Additionally, we present a simplified version of the best-known algorithms that match our impossibility result.
title Tight Bounds for Answering Adaptively Chosen Concentrated Queries
topic Data Structures and Algorithms
Machine Learning
url https://arxiv.org/abs/2507.13700