One Model Many Scores: Using Multiverse Analysis to Prevent Fairness Hacking and Evaluate the Influence of Model Design Decisions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Simson, Jan, Pfisterer, Florian, Kern, Christoph |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Lazy Data Practices Harm Fairness Research
von: Simson, Jan, et al.
Veröffentlicht: (2024)
von: Simson, Jan, et al.
Veröffentlicht: (2024)
Bias Begins with Data: The FairGround Corpus for Robust and Reproducible Research on Algorithmic Fairness
von: Simson, Jan, et al.
Veröffentlicht: (2025)
von: Simson, Jan, et al.
Veröffentlicht: (2025)
Many AI Analysts, One Dataset: Navigating the Agentic Data Science Multiverse
von: Bertran, Martin, et al.
Veröffentlicht: (2026)
von: Bertran, Martin, et al.
Veröffentlicht: (2026)
Connecting Algorithmic Fairness to Quality Dimensions in Machine Learning in Official Statistics and Survey Production
von: Schenk, Patrick Oliver, et al.
Veröffentlicht: (2024)
von: Schenk, Patrick Oliver, et al.
Veröffentlicht: (2024)
On Prediction-Modelers and Decision-Makers: Why Fairness Requires More Than a Fair Prediction Model
von: Scantamburlo, Teresa, et al.
Veröffentlicht: (2023)
von: Scantamburlo, Teresa, et al.
Veröffentlicht: (2023)
The Fairness of Credit Scoring Models
von: Hurlin, Christophe, et al.
Veröffentlicht: (2022)
von: Hurlin, Christophe, et al.
Veröffentlicht: (2022)
Flextron: Many-in-One Flexible Large Language Model
von: Cai, Ruisi, et al.
Veröffentlicht: (2024)
von: Cai, Ruisi, et al.
Veröffentlicht: (2024)
Out of One, Many: Using Language Models to Simulate Human Samples
von: Argyle, Lisa P., et al.
Veröffentlicht: (2022)
von: Argyle, Lisa P., et al.
Veröffentlicht: (2022)
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale
von: Roth, Amit, et al.
Veröffentlicht: (2026)
von: Roth, Amit, et al.
Veröffentlicht: (2026)
The C-index Multiverse
von: Sierra, Begoña B., et al.
Veröffentlicht: (2025)
von: Sierra, Begoña B., et al.
Veröffentlicht: (2025)
EuroSpeech: A Multilingual Speech Corpus
von: Pfisterer, Samuel, et al.
Veröffentlicht: (2025)
von: Pfisterer, Samuel, et al.
Veröffentlicht: (2025)
OneFlowSBI: One Model, Many Queries for Simulation-Based Inference
von: Nautiyal, Mayank, et al.
Veröffentlicht: (2026)
von: Nautiyal, Mayank, et al.
Veröffentlicht: (2026)
Multiverse: Your Language Models Secretly Decide How to Parallelize and Merge Generation
von: Yang, Xinyu, et al.
Veröffentlicht: (2025)
von: Yang, Xinyu, et al.
Veröffentlicht: (2025)
Fairness Hacking: The Malicious Practice of Shrouding Unfairness in Algorithms
von: Meding, Kristof, et al.
Veröffentlicht: (2023)
von: Meding, Kristof, et al.
Veröffentlicht: (2023)
Guiding LLM Decision-Making with Fairness Reward Models
von: Hall, Zara, et al.
Veröffentlicht: (2025)
von: Hall, Zara, et al.
Veröffentlicht: (2025)
Diffusion Attribution Score: Evaluating Training Data Influence in Diffusion Models
von: Lin, Jinxu, et al.
Veröffentlicht: (2024)
von: Lin, Jinxu, et al.
Veröffentlicht: (2024)
Fair Decisions from Calibrated Scores: Achieving Optimal Classification While Satisfying Sufficiency
von: Benger, Etam, et al.
Veröffentlicht: (2026)
von: Benger, Etam, et al.
Veröffentlicht: (2026)
On Benchmark Hacking in ML Contests: Modeling, Insights and Design
von: Qiu, Xiaoyun, et al.
Veröffentlicht: (2026)
von: Qiu, Xiaoyun, et al.
Veröffentlicht: (2026)
Inference-Time Reward Hacking in Large Language Models
von: Khalaf, Hadi, et al.
Veröffentlicht: (2025)
von: Khalaf, Hadi, et al.
Veröffentlicht: (2025)
From Many Models, One: Macroeconomic Forecasting with Reservoir Ensembles
von: Ballarin, Giovanni, et al.
Veröffentlicht: (2025)
von: Ballarin, Giovanni, et al.
Veröffentlicht: (2025)
Mind the Gap: Measuring Generalization Performance Across Multiple Objectives
von: Feurer, Matthias, et al.
Veröffentlicht: (2022)
von: Feurer, Matthias, et al.
Veröffentlicht: (2022)
Mapping the Multiverse of Latent Representations
von: Wayland, Jeremy, et al.
Veröffentlicht: (2024)
von: Wayland, Jeremy, et al.
Veröffentlicht: (2024)
Honesty to Subterfuge: In-Context Reinforcement Learning Can Make Honest Models Reward Hack
von: McKee-Reid, Leo, et al.
Veröffentlicht: (2024)
von: McKee-Reid, Leo, et al.
Veröffentlicht: (2024)
On Teacher Hacking in Language Model Distillation
von: Tiapkin, Daniil, et al.
Veröffentlicht: (2025)
von: Tiapkin, Daniil, et al.
Veröffentlicht: (2025)
Privacy Constrained Fairness Estimation for Decision Trees
von: van der Steen, Florian, et al.
Veröffentlicht: (2023)
von: van der Steen, Florian, et al.
Veröffentlicht: (2023)
A Bias-Variance-Covariance Decomposition of Kernel Scores for Generative Models
von: Gruber, Sebastian G., et al.
Veröffentlicht: (2023)
von: Gruber, Sebastian G., et al.
Veröffentlicht: (2023)
Mitigating Preference Hacking in Policy Optimization with Pessimism
von: Gupta, Dhawal, et al.
Veröffentlicht: (2025)
von: Gupta, Dhawal, et al.
Veröffentlicht: (2025)
One Model, Many Skills: Parameter-Efficient Fine-Tuning for Multitask Code Analysis
von: Akli, Amal, et al.
Veröffentlicht: (2026)
von: Akli, Amal, et al.
Veröffentlicht: (2026)
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
von: Wang, Xiaohua, et al.
Veröffentlicht: (2026)
von: Wang, Xiaohua, et al.
Veröffentlicht: (2026)
One Head, Many Models: Cross-Attention Routing for Cost-Aware LLM Selection
von: Pulishetty, Roshini, et al.
Veröffentlicht: (2025)
von: Pulishetty, Roshini, et al.
Veröffentlicht: (2025)
Hacking Predictors Means Hacking Cars: Using Sensitivity Analysis to Identify Trajectory Prediction Vulnerabilities for Autonomous Driving Security
von: Gibson, Marsalis, et al.
Veröffentlicht: (2024)
von: Gibson, Marsalis, et al.
Veröffentlicht: (2024)
Checkmating One, by Using Many: Combining Mixture of Experts with MCTS to Improve in Chess
von: Helfenstein, Felix, et al.
Veröffentlicht: (2024)
von: Helfenstein, Felix, et al.
Veröffentlicht: (2024)
Evaluating Fairness in Transaction Fraud Models: Fairness Metrics, Bias Audits, and Challenges
von: Kamalaruban, Parameswaran, et al.
Veröffentlicht: (2024)
von: Kamalaruban, Parameswaran, et al.
Veröffentlicht: (2024)
Whence Is A Model Fair? Fixing Fairness Bugs via Propensity Score Matching
von: Peng, Kewen, et al.
Veröffentlicht: (2025)
von: Peng, Kewen, et al.
Veröffentlicht: (2025)
Evaluating Binary Decision Biases in Large Language Models: Implications for Fair Agent-Based Financial Simulations
von: Vidler, Alicia, et al.
Veröffentlicht: (2025)
von: Vidler, Alicia, et al.
Veröffentlicht: (2025)
Evaluating Posterior Probabilities: Decision Theory, Proper Scoring Rules, and Calibration
von: Ferrer, Luciana, et al.
Veröffentlicht: (2024)
von: Ferrer, Luciana, et al.
Veröffentlicht: (2024)
Enhancing Group Fairness in Online Settings Using Oblique Decision Forests
von: Chowdhury, Somnath Basu Roy, et al.
Veröffentlicht: (2023)
von: Chowdhury, Somnath Basu Roy, et al.
Veröffentlicht: (2023)
FairUDT: Fairness-aware Uplift Decision Trees
von: Zahid, Anam, et al.
Veröffentlicht: (2025)
von: Zahid, Anam, et al.
Veröffentlicht: (2025)
Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
von: Eisenstein, Jacob, et al.
Veröffentlicht: (2023)
von: Eisenstein, Jacob, et al.
Veröffentlicht: (2023)
Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking
von: Miao, Yuchun, et al.
Veröffentlicht: (2025)
von: Miao, Yuchun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Lazy Data Practices Harm Fairness Research
von: Simson, Jan, et al.
Veröffentlicht: (2024) -
Bias Begins with Data: The FairGround Corpus for Robust and Reproducible Research on Algorithmic Fairness
von: Simson, Jan, et al.
Veröffentlicht: (2025) -
Many AI Analysts, One Dataset: Navigating the Agentic Data Science Multiverse
von: Bertran, Martin, et al.
Veröffentlicht: (2026) -
Connecting Algorithmic Fairness to Quality Dimensions in Machine Learning in Official Statistics and Survey Production
von: Schenk, Patrick Oliver, et al.
Veröffentlicht: (2024) -
On Prediction-Modelers and Decision-Makers: Why Fairness Requires More Than a Fair Prediction Model
von: Scantamburlo, Teresa, et al.
Veröffentlicht: (2023)