Position: Measure Dataset Diversity, Don't Just Claim It
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Dora, Andrews, Jerone T. A., Papakyriakopoulos, Orestis, Xiang, Alice |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Taxonomy of Challenges to Curating Fair Datasets
by: Zhao, Dora, et al.
Published: (2024)
by: Zhao, Dora, et al.
Published: (2024)
Resampled Datasets Are Not Enough: Mitigating Societal Bias Beyond Single Attributes
by: Hirota, Yusuke, et al.
Published: (2024)
by: Hirota, Yusuke, et al.
Published: (2024)
Why Don't Prompt-Based Fairness Metrics Correlate?
by: Zayed, Abdelrahman, et al.
Published: (2024)
by: Zayed, Abdelrahman, et al.
Published: (2024)
Information Retrieval Induced Safety Degradation in AI Agents
by: Yu, Cheng, et al.
Published: (2025)
by: Yu, Cheng, et al.
Published: (2025)
The Compliance Gap: Why AI Systems Promise to Follow Process Instructions but Don't
by: Shin, Kwan Soo
Published: (2026)
by: Shin, Kwan Soo
Published: (2026)
Reasoning Models Don't Just Think Longer, They Move Differently
by: Gjølbye, Anders, et al.
Published: (2026)
by: Gjølbye, Anders, et al.
Published: (2026)
How Should AI Safety Benchmarks Benchmark Safety?
by: Yu, Cheng, et al.
Published: (2026)
by: Yu, Cheng, et al.
Published: (2026)
Not My Voice! A Taxonomy of Ethical and Safety Harms of Speech Generators
by: Hutiri, Wiebke, et al.
Published: (2024)
by: Hutiri, Wiebke, et al.
Published: (2024)
ClaimVer: Explainable Claim-Level Verification and Evidence Attribution of Text Through Knowledge Graphs
by: Dammu, Preetam Prabhu Srikar, et al.
Published: (2024)
by: Dammu, Preetam Prabhu Srikar, et al.
Published: (2024)
Condense, Don't Just Prune: Enhancing Efficiency and Performance in MoE Layer Pruning
by: Cao, Mingyu, et al.
Published: (2024)
by: Cao, Mingyu, et al.
Published: (2024)
Opportunities of Reinforcement Learning in South Africa's Just Transition
by: Formanek, Claude, et al.
Published: (2024)
by: Formanek, Claude, et al.
Published: (2024)
Don't Just Say "I don't know"! Self-aligning Large Language Models for Responding to Unknown Questions with Explanations
by: Deng, Yang, et al.
Published: (2024)
by: Deng, Yang, et al.
Published: (2024)
Efficient Bias Mitigation Without Privileged Information
by: Zarlenga, Mateo Espinosa, et al.
Published: (2024)
by: Zarlenga, Mateo Espinosa, et al.
Published: (2024)
If You Don't Understand It, Don't Use It: Eliminating Trojans with Filters Between Layers
by: Hernandez, Adriano
Published: (2024)
by: Hernandez, Adriano
Published: (2024)
JustQ: Automated Deployment of Fair and Accurate Quantum Neural Networks
by: Wang, Ruhan, et al.
Published: (2024)
by: Wang, Ruhan, et al.
Published: (2024)
The Fragility of Fairness: Causal Sensitivity Analysis for Fair Machine Learning
by: Fawkes, Jake, et al.
Published: (2024)
by: Fawkes, Jake, et al.
Published: (2024)
Humanizing LLMs: A Survey of Psychological Measurements with Tools, Datasets, and Human-Agent Applications
by: Dong, Wenhan, et al.
Published: (2025)
by: Dong, Wenhan, et al.
Published: (2025)
Rethinking Anonymity Claims in Synthetic Data Generation: A Model-Centric Privacy Attack Perspective
by: Ganev, Georgi, et al.
Published: (2026)
by: Ganev, Georgi, et al.
Published: (2026)
Quality-Diversity Generative Sampling for Learning with Synthetic Data
by: Chang, Allen, et al.
Published: (2023)
by: Chang, Allen, et al.
Published: (2023)
Don't Be Greedy, Just Relax! Pruning LLMs via Frank-Wolfe
by: Roux, Christophe, et al.
Published: (2025)
by: Roux, Christophe, et al.
Published: (2025)
Dataset Representativeness and Downstream Task Fairness
by: Borza, Victor, et al.
Published: (2024)
by: Borza, Victor, et al.
Published: (2024)
Rehabilitating Homeless: Dataset and Key Insights
by: Bykova, Anna, et al.
Published: (2023)
by: Bykova, Anna, et al.
Published: (2023)
Earth Embeddings Reveal Diverse Urban Signals from Space
by: Gong, Wenjing, et al.
Published: (2026)
by: Gong, Wenjing, et al.
Published: (2026)
Computational Measurement of Political Positions: A Review of Text-Based Ideal Point Estimation Algorithms
by: Parschan, Patrick, et al.
Published: (2025)
by: Parschan, Patrick, et al.
Published: (2025)
Position: Don't be Afraid of Over-Smoothing And Over-Squashing
by: Kormann, Niklas, et al.
Published: (2026)
by: Kormann, Niklas, et al.
Published: (2026)
Long-Term Fairness in Sequential Multi-Agent Selection with Positive Reinforcement
by: Puranik, Bhagyashree, et al.
Published: (2024)
by: Puranik, Bhagyashree, et al.
Published: (2024)
Uncertainty-based Fairness Measures
by: Kuzucu, Selim, et al.
Published: (2023)
by: Kuzucu, Selim, et al.
Published: (2023)
Evaluating the Performance of Large Language Models in Scientific Claim Detection and Classification
by: Faruk, Tanjim Bin
Published: (2024)
by: Faruk, Tanjim Bin
Published: (2024)
A Water Efficiency Dataset for African Data Centers
by: Shumba, Noah, et al.
Published: (2024)
by: Shumba, Noah, et al.
Published: (2024)
Position: Stop Preaching and Start Practising Data Frugality for Responsible Development of AI
by: Wilson, Sophia N., et al.
Published: (2026)
by: Wilson, Sophia N., et al.
Published: (2026)
How Robust is your Fair Model? Exploring the Robustness of Diverse Fairness Strategies
by: Small, Edward, et al.
Published: (2022)
by: Small, Edward, et al.
Published: (2022)
FairHome: A Fair Housing and Fair Lending Dataset
by: Bagalkotkar, Anusha, et al.
Published: (2024)
by: Bagalkotkar, Anusha, et al.
Published: (2024)
Fair for a few: Improving Fairness in Doubly Imbalanced Datasets
by: Yalcin, Ata, et al.
Published: (2025)
by: Yalcin, Ata, et al.
Published: (2025)
Counterfactual Fairness Evaluation of Machine Learning Models on Educational Datasets
by: Kim, Woojin, et al.
Published: (2025)
by: Kim, Woojin, et al.
Published: (2025)
DCAST: Diverse Class-Aware Self-Training Mitigates Selection Bias for Fairer Learning
by: Tepeli, Yasin I., et al.
Published: (2024)
by: Tepeli, Yasin I., et al.
Published: (2024)
From Model Performance to Claim: How a Change of Focus in Machine Learning Replicability Can Help Bridge the Responsibility Gap
by: Kou, Tianqi
Published: (2024)
by: Kou, Tianqi
Published: (2024)
Measuring and Mitigating Biases in Motor Insurance Pricing
by: Moriah, Mulah, et al.
Published: (2023)
by: Moriah, Mulah, et al.
Published: (2023)
Addressing Shortcomings in Fair Graph Learning Datasets: Towards a New Benchmark
by: Qian, Xiaowei, et al.
Published: (2024)
by: Qian, Xiaowei, et al.
Published: (2024)
Adaptive Recruitment Resource Allocation to Improve Cohort Representativeness in Participatory Biomedical Datasets
by: Borza, Victor, et al.
Published: (2024)
by: Borza, Victor, et al.
Published: (2024)
A Primer on Causal and Statistical Dataset Biases for Fair and Robust Image Analysis
by: Jones, Charles, et al.
Published: (2025)
by: Jones, Charles, et al.
Published: (2025)
Similar Items
-
A Taxonomy of Challenges to Curating Fair Datasets
by: Zhao, Dora, et al.
Published: (2024) -
Resampled Datasets Are Not Enough: Mitigating Societal Bias Beyond Single Attributes
by: Hirota, Yusuke, et al.
Published: (2024) -
Why Don't Prompt-Based Fairness Metrics Correlate?
by: Zayed, Abdelrahman, et al.
Published: (2024) -
Information Retrieval Induced Safety Degradation in AI Agents
by: Yu, Cheng, et al.
Published: (2025) -
The Compliance Gap: Why AI Systems Promise to Follow Process Instructions but Don't
by: Shin, Kwan Soo
Published: (2026)