Metrics for Inter-Dataset Similarity with Example Applications in Synthetic Data and Feature Selection Evaluation -- Extended Version
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rajabinasab, Muhammad, Lautrup, Anton D., Zimek, Arthur |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FSDEM: Feature Selection Dynamic Evaluation Metric
von: Rajabinasab, Muhammad, et al.
Veröffentlicht: (2024)
von: Rajabinasab, Muhammad, et al.
Veröffentlicht: (2024)
FSEVAL: Feature Selection Evaluation Toolbox and Dashboard
von: Rajabinasab, Muhammad, et al.
Veröffentlicht: (2026)
von: Rajabinasab, Muhammad, et al.
Veröffentlicht: (2026)
Disjoint Generative Models
von: Lautrup, Anton Danholt, et al.
Veröffentlicht: (2025)
von: Lautrup, Anton Danholt, et al.
Veröffentlicht: (2025)
SynthEval: A Framework for Detailed Utility and Privacy Evaluation of Tabular Synthetic Data
von: Lautrup, Anton Danholt, et al.
Veröffentlicht: (2024)
von: Lautrup, Anton Danholt, et al.
Veröffentlicht: (2024)
Worse than Random: The Importance of a Baseline for Unsupervised Feature Selection
von: Rajabinasab, Muhammad, et al.
Veröffentlicht: (2026)
von: Rajabinasab, Muhammad, et al.
Veröffentlicht: (2026)
Sharing is CAIRing: Characterizing Principles and Assessing Properties of Universal Privacy Evaluation for Synthetic Tabular Data
von: Hyrup, Tobias, et al.
Veröffentlicht: (2023)
von: Hyrup, Tobias, et al.
Veröffentlicht: (2023)
MARS: Magnitude-Aware Rank Statistics
von: Rajabinasab, Muhammad, et al.
Veröffentlicht: (2026)
von: Rajabinasab, Muhammad, et al.
Veröffentlicht: (2026)
Randomized PCA Forest for Unsupervised Outlier Detection
von: Rajabinasab, Muhammad, et al.
Veröffentlicht: (2025)
von: Rajabinasab, Muhammad, et al.
Veröffentlicht: (2025)
Robust Statistical Scaling of Outlier Scores: Improving the Quality of Outlier Probabilities for Outliers (Extended Version)
von: Röchner, Philipp, et al.
Veröffentlicht: (2024)
von: Röchner, Philipp, et al.
Veröffentlicht: (2024)
Analyzing Similarity Metrics for Data Selection for Language Model Pretraining
von: Sam, Dylan, et al.
Veröffentlicht: (2025)
von: Sam, Dylan, et al.
Veröffentlicht: (2025)
Evaluating Inter-Column Logical Relationships in Synthetic Tabular Data Generation
von: Long, Yunbo, et al.
Veröffentlicht: (2025)
von: Long, Yunbo, et al.
Veröffentlicht: (2025)
Transparent Neighborhood Approximation for Text Classifier Explanation
von: Cai, Yi, et al.
Veröffentlicht: (2024)
von: Cai, Yi, et al.
Veröffentlicht: (2024)
Achieving Hilbert-Schmidt Independence Under Rényi Differential Privacy for Fair and Private Data Generation
von: Hyrup, Tobias, et al.
Veröffentlicht: (2025)
von: Hyrup, Tobias, et al.
Veröffentlicht: (2025)
The Inadequacy of Similarity-based Privacy Metrics: Privacy Attacks against "Truly Anonymous" Synthetic Datasets
von: Ganev, Georgi, et al.
Veröffentlicht: (2023)
von: Ganev, Georgi, et al.
Veröffentlicht: (2023)
DCFO: Density-Based Counterfactuals for Outliers -- Additional Material
von: Amico, Tommaso, et al.
Veröffentlicht: (2025)
von: Amico, Tommaso, et al.
Veröffentlicht: (2025)
SDQM: Synthetic Data Quality Metric for Object Detection Dataset Evaluation
von: Zenith, Ayush, et al.
Veröffentlicht: (2025)
von: Zenith, Ayush, et al.
Veröffentlicht: (2025)
Designing Informative Metrics for Few-Shot Example Selection
von: Adiga, Rishabh, et al.
Veröffentlicht: (2024)
von: Adiga, Rishabh, et al.
Veröffentlicht: (2024)
A Universal Metric of Dataset Similarity for Cross-silo Federated Learning
von: Elhussein, Ahmed, et al.
Veröffentlicht: (2024)
von: Elhussein, Ahmed, et al.
Veröffentlicht: (2024)
Memisis: Orchestrating and Evaluating Synthetic Data for Tabular Health Datasets
von: Nagesh, Nitish, et al.
Veröffentlicht: (2026)
von: Nagesh, Nitish, et al.
Veröffentlicht: (2026)
Synthetic Data Privacy Metrics
von: Steier, Amy, et al.
Veröffentlicht: (2025)
von: Steier, Amy, et al.
Veröffentlicht: (2025)
PEAKS: Selecting Key Training Examples Incrementally via Prediction Error Anchored by Kernel Similarity
von: Gurbuz, Mustafa Burak, et al.
Veröffentlicht: (2025)
von: Gurbuz, Mustafa Burak, et al.
Veröffentlicht: (2025)
SynQuE: Estimating Synthetic Dataset Quality Without Annotations
von: Chen, Arthur, et al.
Veröffentlicht: (2025)
von: Chen, Arthur, et al.
Veröffentlicht: (2025)
TabSynDex: A Universal Metric for Robust Evaluation of Synthetic Tabular Data
von: Chundawat, Vikram S, et al.
Veröffentlicht: (2022)
von: Chundawat, Vikram S, et al.
Veröffentlicht: (2022)
Evaluating Synthetic Tabular Data Generated To Augment Small Sample Datasets
von: Marin, Javier
Veröffentlicht: (2022)
von: Marin, Javier
Veröffentlicht: (2022)
ICAFS: Inter-Client-Aware Feature Selection for Vertical Federated Learning
von: Jin, Ruochen, et al.
Veröffentlicht: (2025)
von: Jin, Ruochen, et al.
Veröffentlicht: (2025)
The LLM Data Auditor: A Metric-oriented Survey on Quality and Trustworthiness in Evaluating Synthetic Data
von: Zhang, Kaituo, et al.
Veröffentlicht: (2026)
von: Zhang, Kaituo, et al.
Veröffentlicht: (2026)
Metric Design != Metric Behavior: Improving Metric Selection for the Unbiased Evaluation of Dimensionality Reduction
von: Bae, Jiyeon, et al.
Veröffentlicht: (2025)
von: Bae, Jiyeon, et al.
Veröffentlicht: (2025)
Automatic Unsupervised Ensemble Outlier Model Selection--Extended Version
von: Phan, Hong-Phuc, et al.
Veröffentlicht: (2026)
von: Phan, Hong-Phuc, et al.
Veröffentlicht: (2026)
SynQP: A Framework and Metrics for Evaluating the Quality and Privacy Risk of Synthetic Data
von: Hu, Bing, et al.
Veröffentlicht: (2026)
von: Hu, Bing, et al.
Veröffentlicht: (2026)
Prioritizing Informative Features and Examples for Deep Learning from Noisy Data
von: Park, Dongmin
Veröffentlicht: (2024)
von: Park, Dongmin
Veröffentlicht: (2024)
It's My Data Too: Private ML for Datasets with Multi-User Training Examples
von: Ganesh, Arun, et al.
Veröffentlicht: (2025)
von: Ganesh, Arun, et al.
Veröffentlicht: (2025)
Variance-Adjusted Cosine Distance as Similarity Metric
von: Sahoo, Satyajeet, et al.
Veröffentlicht: (2025)
von: Sahoo, Satyajeet, et al.
Veröffentlicht: (2025)
High-dimensional Analysis of Synthetic Data Selection
von: Rezaei, Parham, et al.
Veröffentlicht: (2025)
von: Rezaei, Parham, et al.
Veröffentlicht: (2025)
Structure-aware Hybrid-order Similarity Learning for Multi-view Unsupervised Feature Selection
von: Xu, Lin, et al.
Veröffentlicht: (2025)
von: Xu, Lin, et al.
Veröffentlicht: (2025)
Universal Feature Selection for Simultaneous Interpretability of Multitask Datasets
von: Raymond, Matt, et al.
Veröffentlicht: (2024)
von: Raymond, Matt, et al.
Veröffentlicht: (2024)
On the (In)Significance of Feature Selection in High-Dimensional Datasets
von: Neekhra, Bhavesh, et al.
Veröffentlicht: (2025)
von: Neekhra, Bhavesh, et al.
Veröffentlicht: (2025)
Reducing Instability in Synthetic Data Evaluation with a Super-Metric in MalDataGen
von: da Silva, Anna Luiza Gomes, et al.
Veröffentlicht: (2025)
von: da Silva, Anna Luiza Gomes, et al.
Veröffentlicht: (2025)
SHAP Distance: An Explainability-Aware Metric for Evaluating the Semantic Fidelity of Synthetic Tabular Data
von: Yu, Ke, et al.
Veröffentlicht: (2025)
von: Yu, Ke, et al.
Veröffentlicht: (2025)
Mechanisms for Data Sharing in Collaborative Causal Inference (Extended Version)
von: Filter, Björn, et al.
Veröffentlicht: (2024)
von: Filter, Björn, et al.
Veröffentlicht: (2024)
QCore: Data-Efficient, On-Device Continual Calibration for Quantized Models -- Extended Version
von: Campos, David, et al.
Veröffentlicht: (2024)
von: Campos, David, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
FSDEM: Feature Selection Dynamic Evaluation Metric
von: Rajabinasab, Muhammad, et al.
Veröffentlicht: (2024) -
FSEVAL: Feature Selection Evaluation Toolbox and Dashboard
von: Rajabinasab, Muhammad, et al.
Veröffentlicht: (2026) -
Disjoint Generative Models
von: Lautrup, Anton Danholt, et al.
Veröffentlicht: (2025) -
SynthEval: A Framework for Detailed Utility and Privacy Evaluation of Tabular Synthetic Data
von: Lautrup, Anton Danholt, et al.
Veröffentlicht: (2024) -
Worse than Random: The Importance of a Baseline for Unsupervised Feature Selection
von: Rajabinasab, Muhammad, et al.
Veröffentlicht: (2026)