UTOPIA: Unlearnable Tabular Data via Decoupled Shortcut Embedding

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: He, Jiaming, Luo, Fuming, Li, Hongwei, Jiang, Wenbo, Fan, Wenshu, Shi, Zhenbo, Jiang, Xudong, Yu, Yi
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917262500298752
author He, Jiaming
Luo, Fuming
Li, Hongwei
Jiang, Wenbo
Fan, Wenshu
Shi, Zhenbo
Jiang, Xudong
Yu, Yi
author_facet He, Jiaming
Luo, Fuming
Li, Hongwei
Jiang, Wenbo
Fan, Wenshu
Shi, Zhenbo
Jiang, Xudong
Yu, Yi
contents Unlearnable examples (UE) have emerged as a practical mechanism to prevent unauthorized model training on private vision data, while extending this protection to tabular data is nontrivial. Tabular data in finance and healthcare is highly sensitive, yet existing UE methods transfer poorly because tabular features mix numerical and categorical constraints and exhibit saliency sparsity, with learning dominated by a few dimensions. Under a Spectral Dominance condition, we show certified unlearnability is feasible when the poison spectrum overwhelms the clean semantic spectrum. Guided by this, we propose Unlearnable Tabular Data via DecOuPled Shortcut EmbeddIng (UTOPIA), which exploits feature redundancy to decouple optimization into two channels: high saliency features for semantic obfuscation and low saliency redundant features for embedding a hyper correlated shortcut, yielding constraint-aware dominant shortcuts while preserving tabular validity. Extensive experiments across tabular datasets and models show UTOPIA drives unauthorized training toward near random performance, outperforming strong UE baselines and transferring well across architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2602_07358
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle UTOPIA: Unlearnable Tabular Data via Decoupled Shortcut Embedding
He, Jiaming
Luo, Fuming
Li, Hongwei
Jiang, Wenbo
Fan, Wenshu
Shi, Zhenbo
Jiang, Xudong
Yu, Yi
Machine Learning
Unlearnable examples (UE) have emerged as a practical mechanism to prevent unauthorized model training on private vision data, while extending this protection to tabular data is nontrivial. Tabular data in finance and healthcare is highly sensitive, yet existing UE methods transfer poorly because tabular features mix numerical and categorical constraints and exhibit saliency sparsity, with learning dominated by a few dimensions. Under a Spectral Dominance condition, we show certified unlearnability is feasible when the poison spectrum overwhelms the clean semantic spectrum. Guided by this, we propose Unlearnable Tabular Data via DecOuPled Shortcut EmbeddIng (UTOPIA), which exploits feature redundancy to decouple optimization into two channels: high saliency features for semantic obfuscation and low saliency redundant features for embedding a hyper correlated shortcut, yielding constraint-aware dominant shortcuts while preserving tabular validity. Extensive experiments across tabular datasets and models show UTOPIA drives unauthorized training toward near random performance, outperforming strong UE baselines and transferring well across architectures.
title UTOPIA: Unlearnable Tabular Data via Decoupled Shortcut Embedding
topic Machine Learning
url https://arxiv.org/abs/2602.07358