Incentivizing Truthfulness and Collaborative Fairness in Bayesian Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sim, Rachael Hwee Ling, Fan, Jue, Tian, Xiao, Xu, Xinyi, Jaillet, Patrick, Low, Bryan Kian Hsiang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913116612198400
author Sim, Rachael Hwee Ling
Fan, Jue
Tian, Xiao
Xu, Xinyi
Jaillet, Patrick
Low, Bryan Kian Hsiang
author_facet Sim, Rachael Hwee Ling
Fan, Jue
Tian, Xiao
Xu, Xinyi
Jaillet, Patrick
Low, Bryan Kian Hsiang
contents Collaborative machine learning involves training high-quality models using datasets from a number of sources. To incentivize sources to share data, existing data valuation methods fairly reward each source based on its data submitted as is. However, as these methods do not verify nor incentivize data truthfulness, the sources can manipulate their data (e.g., by submitting duplicated or noisy data) to artificially increase their valuations and rewards or prevent others from benefiting. This paper presents the first mechanism that provably ensures (F) collaborative fairness and incentivizes (T) truthfulness at equilibrium for Bayesian models. Our mechanism combines semivalues (e.g., Shapley value), which ensure fairness, and a truthful data valuation function (DVF) based on a validation set that is unknown to the sources. As semivalues are influenced by others' data, we introduce an additional condition to prove that a source can maximize its expected data values in coalitions and semivalues by submitting a dataset that captures its true knowledge. Additionally, we discuss the implications and suitable relaxations of (F) and (T) when the mediator has a limited budget for rewards or lacks a validation set. Our theoretical findings are validated on synthetic and real-world datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2605_11889
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Incentivizing Truthfulness and Collaborative Fairness in Bayesian Learning
Sim, Rachael Hwee Ling
Fan, Jue
Tian, Xiao
Xu, Xinyi
Jaillet, Patrick
Low, Bryan Kian Hsiang
Machine Learning
Artificial Intelligence
Collaborative machine learning involves training high-quality models using datasets from a number of sources. To incentivize sources to share data, existing data valuation methods fairly reward each source based on its data submitted as is. However, as these methods do not verify nor incentivize data truthfulness, the sources can manipulate their data (e.g., by submitting duplicated or noisy data) to artificially increase their valuations and rewards or prevent others from benefiting. This paper presents the first mechanism that provably ensures (F) collaborative fairness and incentivizes (T) truthfulness at equilibrium for Bayesian models. Our mechanism combines semivalues (e.g., Shapley value), which ensure fairness, and a truthful data valuation function (DVF) based on a validation set that is unknown to the sources. As semivalues are influenced by others' data, we introduce an additional condition to prove that a source can maximize its expected data values in coalitions and semivalues by submitting a dataset that captures its true knowledge. Additionally, we discuss the implications and suitable relaxations of (F) and (T) when the mediator has a limited budget for rewards or lacks a validation set. Our theoretical findings are validated on synthetic and real-world datasets.
title Incentivizing Truthfulness and Collaborative Fairness in Bayesian Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.11889