Is Data Shapley Not Better than Random in Data Selection? Ask NASH

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tian, Xiao, Fan, Jue, Sim, Rachael Hwee Ling, Wang, Zixuan, Chen, Nancy F., Low, Bryan Kian Hsiang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909035064721408
author Tian, Xiao
Fan, Jue
Sim, Rachael Hwee Ling
Wang, Zixuan
Chen, Nancy F.
Low, Bryan Kian Hsiang
author_facet Tian, Xiao
Fan, Jue
Sim, Rachael Hwee Ling
Wang, Zixuan
Chen, Nancy F.
Low, Bryan Kian Hsiang
contents Data selection studies the problem of identifying high-quality subsets of training data. While some existing works have considered selecting the subset of data with top-$m$ Data Shapley or other semivalues as they account for the interaction among every subset of data, other works argue that Data Shapley can sometimes perform ineffectively in practice and select subsets that are no better than random. This raises the questions: (I) Are there certain "Shapley-informative" settings where Data Shapley consistently works well? (II) Can we strategically utilize these settings to select high-quality subsets consistently and efficiently? In this paper, we propose a novel data selection framework, NASH (Non-linear Aggregation of SHapley-informative components), which (I) decomposes the target utility function (e.g., validation accuracy) into simpler, Shapley-informative component functions, and selects data by optimizing an objective that (II) aggregates these components non-linearly. We demonstrate that NASH substantially boosts the effectiveness of Shapley/semivalue-based data selection with minimal additional runtime cost.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10684
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Is Data Shapley Not Better than Random in Data Selection? Ask NASH
Tian, Xiao
Fan, Jue
Sim, Rachael Hwee Ling
Wang, Zixuan
Chen, Nancy F.
Low, Bryan Kian Hsiang
Machine Learning
Artificial Intelligence
Data selection studies the problem of identifying high-quality subsets of training data. While some existing works have considered selecting the subset of data with top-$m$ Data Shapley or other semivalues as they account for the interaction among every subset of data, other works argue that Data Shapley can sometimes perform ineffectively in practice and select subsets that are no better than random. This raises the questions: (I) Are there certain "Shapley-informative" settings where Data Shapley consistently works well? (II) Can we strategically utilize these settings to select high-quality subsets consistently and efficiently? In this paper, we propose a novel data selection framework, NASH (Non-linear Aggregation of SHapley-informative components), which (I) decomposes the target utility function (e.g., validation accuracy) into simpler, Shapley-informative component functions, and selects data by optimizing an objective that (II) aggregates these components non-linearly. We demonstrate that NASH substantially boosts the effectiveness of Shapley/semivalue-based data selection with minimal additional runtime cost.
title Is Data Shapley Not Better than Random in Data Selection? Ask NASH
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.10684