Metrics for Inter-Dataset Similarity with Example Applications in Synthetic Data and Feature Selection Evaluation -- Extended Version
Fuente:
arXiv
Saved in:
| Main Authors: | Rajabinasab, Muhammad, Lautrup, Anton D., Zimek, Arthur |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FSDEM: Feature Selection Dynamic Evaluation Metric
by: Rajabinasab, Muhammad, et al.
Published: (2024)
by: Rajabinasab, Muhammad, et al.
Published: (2024)
FSEVAL: Feature Selection Evaluation Toolbox and Dashboard
by: Rajabinasab, Muhammad, et al.
Published: (2026)
by: Rajabinasab, Muhammad, et al.
Published: (2026)
Disjoint Generative Models
by: Lautrup, Anton Danholt, et al.
Published: (2025)
by: Lautrup, Anton Danholt, et al.
Published: (2025)
SynthEval: A Framework for Detailed Utility and Privacy Evaluation of Tabular Synthetic Data
by: Lautrup, Anton Danholt, et al.
Published: (2024)
by: Lautrup, Anton Danholt, et al.
Published: (2024)
Worse than Random: The Importance of a Baseline for Unsupervised Feature Selection
by: Rajabinasab, Muhammad, et al.
Published: (2026)
by: Rajabinasab, Muhammad, et al.
Published: (2026)
Sharing is CAIRing: Characterizing Principles and Assessing Properties of Universal Privacy Evaluation for Synthetic Tabular Data
by: Hyrup, Tobias, et al.
Published: (2023)
by: Hyrup, Tobias, et al.
Published: (2023)
MARS: Magnitude-Aware Rank Statistics
by: Rajabinasab, Muhammad, et al.
Published: (2026)
by: Rajabinasab, Muhammad, et al.
Published: (2026)
Randomized PCA Forest for Unsupervised Outlier Detection
by: Rajabinasab, Muhammad, et al.
Published: (2025)
by: Rajabinasab, Muhammad, et al.
Published: (2025)
Robust Statistical Scaling of Outlier Scores: Improving the Quality of Outlier Probabilities for Outliers (Extended Version)
by: Röchner, Philipp, et al.
Published: (2024)
by: Röchner, Philipp, et al.
Published: (2024)
Analyzing Similarity Metrics for Data Selection for Language Model Pretraining
by: Sam, Dylan, et al.
Published: (2025)
by: Sam, Dylan, et al.
Published: (2025)
Evaluating Inter-Column Logical Relationships in Synthetic Tabular Data Generation
by: Long, Yunbo, et al.
Published: (2025)
by: Long, Yunbo, et al.
Published: (2025)
Transparent Neighborhood Approximation for Text Classifier Explanation
by: Cai, Yi, et al.
Published: (2024)
by: Cai, Yi, et al.
Published: (2024)
Achieving Hilbert-Schmidt Independence Under Rényi Differential Privacy for Fair and Private Data Generation
by: Hyrup, Tobias, et al.
Published: (2025)
by: Hyrup, Tobias, et al.
Published: (2025)
The Inadequacy of Similarity-based Privacy Metrics: Privacy Attacks against "Truly Anonymous" Synthetic Datasets
by: Ganev, Georgi, et al.
Published: (2023)
by: Ganev, Georgi, et al.
Published: (2023)
DCFO: Density-Based Counterfactuals for Outliers -- Additional Material
by: Amico, Tommaso, et al.
Published: (2025)
by: Amico, Tommaso, et al.
Published: (2025)
SDQM: Synthetic Data Quality Metric for Object Detection Dataset Evaluation
by: Zenith, Ayush, et al.
Published: (2025)
by: Zenith, Ayush, et al.
Published: (2025)
Designing Informative Metrics for Few-Shot Example Selection
by: Adiga, Rishabh, et al.
Published: (2024)
by: Adiga, Rishabh, et al.
Published: (2024)
A Universal Metric of Dataset Similarity for Cross-silo Federated Learning
by: Elhussein, Ahmed, et al.
Published: (2024)
by: Elhussein, Ahmed, et al.
Published: (2024)
Memisis: Orchestrating and Evaluating Synthetic Data for Tabular Health Datasets
by: Nagesh, Nitish, et al.
Published: (2026)
by: Nagesh, Nitish, et al.
Published: (2026)
Synthetic Data Privacy Metrics
by: Steier, Amy, et al.
Published: (2025)
by: Steier, Amy, et al.
Published: (2025)
PEAKS: Selecting Key Training Examples Incrementally via Prediction Error Anchored by Kernel Similarity
by: Gurbuz, Mustafa Burak, et al.
Published: (2025)
by: Gurbuz, Mustafa Burak, et al.
Published: (2025)
SynQuE: Estimating Synthetic Dataset Quality Without Annotations
by: Chen, Arthur, et al.
Published: (2025)
by: Chen, Arthur, et al.
Published: (2025)
TabSynDex: A Universal Metric for Robust Evaluation of Synthetic Tabular Data
by: Chundawat, Vikram S, et al.
Published: (2022)
by: Chundawat, Vikram S, et al.
Published: (2022)
Evaluating Synthetic Tabular Data Generated To Augment Small Sample Datasets
by: Marin, Javier
Published: (2022)
by: Marin, Javier
Published: (2022)
ICAFS: Inter-Client-Aware Feature Selection for Vertical Federated Learning
by: Jin, Ruochen, et al.
Published: (2025)
by: Jin, Ruochen, et al.
Published: (2025)
The LLM Data Auditor: A Metric-oriented Survey on Quality and Trustworthiness in Evaluating Synthetic Data
by: Zhang, Kaituo, et al.
Published: (2026)
by: Zhang, Kaituo, et al.
Published: (2026)
Metric Design != Metric Behavior: Improving Metric Selection for the Unbiased Evaluation of Dimensionality Reduction
by: Bae, Jiyeon, et al.
Published: (2025)
by: Bae, Jiyeon, et al.
Published: (2025)
Automatic Unsupervised Ensemble Outlier Model Selection--Extended Version
by: Phan, Hong-Phuc, et al.
Published: (2026)
by: Phan, Hong-Phuc, et al.
Published: (2026)
SynQP: A Framework and Metrics for Evaluating the Quality and Privacy Risk of Synthetic Data
by: Hu, Bing, et al.
Published: (2026)
by: Hu, Bing, et al.
Published: (2026)
Prioritizing Informative Features and Examples for Deep Learning from Noisy Data
by: Park, Dongmin
Published: (2024)
by: Park, Dongmin
Published: (2024)
It's My Data Too: Private ML for Datasets with Multi-User Training Examples
by: Ganesh, Arun, et al.
Published: (2025)
by: Ganesh, Arun, et al.
Published: (2025)
Variance-Adjusted Cosine Distance as Similarity Metric
by: Sahoo, Satyajeet, et al.
Published: (2025)
by: Sahoo, Satyajeet, et al.
Published: (2025)
High-dimensional Analysis of Synthetic Data Selection
by: Rezaei, Parham, et al.
Published: (2025)
by: Rezaei, Parham, et al.
Published: (2025)
Structure-aware Hybrid-order Similarity Learning for Multi-view Unsupervised Feature Selection
by: Xu, Lin, et al.
Published: (2025)
by: Xu, Lin, et al.
Published: (2025)
Universal Feature Selection for Simultaneous Interpretability of Multitask Datasets
by: Raymond, Matt, et al.
Published: (2024)
by: Raymond, Matt, et al.
Published: (2024)
On the (In)Significance of Feature Selection in High-Dimensional Datasets
by: Neekhra, Bhavesh, et al.
Published: (2025)
by: Neekhra, Bhavesh, et al.
Published: (2025)
Reducing Instability in Synthetic Data Evaluation with a Super-Metric in MalDataGen
by: da Silva, Anna Luiza Gomes, et al.
Published: (2025)
by: da Silva, Anna Luiza Gomes, et al.
Published: (2025)
SHAP Distance: An Explainability-Aware Metric for Evaluating the Semantic Fidelity of Synthetic Tabular Data
by: Yu, Ke, et al.
Published: (2025)
by: Yu, Ke, et al.
Published: (2025)
Mechanisms for Data Sharing in Collaborative Causal Inference (Extended Version)
by: Filter, Björn, et al.
Published: (2024)
by: Filter, Björn, et al.
Published: (2024)
QCore: Data-Efficient, On-Device Continual Calibration for Quantized Models -- Extended Version
by: Campos, David, et al.
Published: (2024)
by: Campos, David, et al.
Published: (2024)
Similar Items
-
FSDEM: Feature Selection Dynamic Evaluation Metric
by: Rajabinasab, Muhammad, et al.
Published: (2024) -
FSEVAL: Feature Selection Evaluation Toolbox and Dashboard
by: Rajabinasab, Muhammad, et al.
Published: (2026) -
Disjoint Generative Models
by: Lautrup, Anton Danholt, et al.
Published: (2025) -
SynthEval: A Framework for Detailed Utility and Privacy Evaluation of Tabular Synthetic Data
by: Lautrup, Anton Danholt, et al.
Published: (2024) -
Worse than Random: The Importance of a Baseline for Unsupervised Feature Selection
by: Rajabinasab, Muhammad, et al.
Published: (2026)