Salvato in:
Dettagli Bibliografici
Autori principali: Huang, Jimin, Feng, Duanyu, Chen, Nuo, Wang, Xiaoyu, Zhang, Zhiqiang, Peng, Xueqing, Lin, Mingquan, Tiwari, Prayag, Xiong, Guojun, Lopez-Lira, Alejandro, Ananiadou, Sophia
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:https://arxiv.org/abs/2605.09855
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911691579588608
author Huang, Jimin
Feng, Duanyu
Chen, Nuo
Wang, Xiaoyu
Zhang, Zhiqiang
Peng, Xueqing
Lin, Mingquan
Tiwari, Prayag
Xiong, Guojun
Lopez-Lira, Alejandro
Ananiadou, Sophia
author_facet Huang, Jimin
Feng, Duanyu
Chen, Nuo
Wang, Xiaoyu
Zhang, Zhiqiang
Peng, Xueqing
Lin, Mingquan
Tiwari, Prayag
Xiong, Guojun
Lopez-Lira, Alejandro
Ananiadou, Sophia
contents Federated learning (FL) enables training large language models (LLMs) without sharing raw data, but adapting LLMs under strict data isolation and non-IID client distributions remains challenging in practice. Synthetic data offers a natural privacy-preserving surrogate for local training, yet existing federated pipelines typically treat synthetic generation as static or loosely coupled with downstream optimization, leading to rapidly diminishing utility under heterogeneous clients. We study federated adaptation of LLMs on tabular tasks where raw records and validation data cannot be shared, and local training must rely entirely on synthetic tables. We propose Concordia, a tri-level optimization framework that aligns synthetic data generation with federated validation utility despite these constraints. At the client level, models are adapted via parameter-efficient LoRA training on synthetic tables. Clients additionally learn lightweight utility scorers from private validation feedback to reweight synthetic samples during local training. At the outer level, each client refines its own synthetic table generator using group-relative policy optimization (GRPO), guided by an ensemble of heterogeneous scorers shared across clients, without aggregating generator parameters or exposing validation data. Experiments on privacy-sensitive tabular benchmarks from finance and healthcare demonstrate that Concordia consistently improves federated performance, cross-client stability, and robustness to distribution shift compared to static and decoupled synthetic-data baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2605_09855
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Concordia: Self-Improving Synthetic Tables for Federated LLMs
Huang, Jimin
Feng, Duanyu
Chen, Nuo
Wang, Xiaoyu
Zhang, Zhiqiang
Peng, Xueqing
Lin, Mingquan
Tiwari, Prayag
Xiong, Guojun
Lopez-Lira, Alejandro
Ananiadou, Sophia
Machine Learning
Federated learning (FL) enables training large language models (LLMs) without sharing raw data, but adapting LLMs under strict data isolation and non-IID client distributions remains challenging in practice. Synthetic data offers a natural privacy-preserving surrogate for local training, yet existing federated pipelines typically treat synthetic generation as static or loosely coupled with downstream optimization, leading to rapidly diminishing utility under heterogeneous clients. We study federated adaptation of LLMs on tabular tasks where raw records and validation data cannot be shared, and local training must rely entirely on synthetic tables. We propose Concordia, a tri-level optimization framework that aligns synthetic data generation with federated validation utility despite these constraints. At the client level, models are adapted via parameter-efficient LoRA training on synthetic tables. Clients additionally learn lightweight utility scorers from private validation feedback to reweight synthetic samples during local training. At the outer level, each client refines its own synthetic table generator using group-relative policy optimization (GRPO), guided by an ensemble of heterogeneous scorers shared across clients, without aggregating generator parameters or exposing validation data. Experiments on privacy-sensitive tabular benchmarks from finance and healthcare demonstrate that Concordia consistently improves federated performance, cross-client stability, and robustness to distribution shift compared to static and decoupled synthetic-data baselines.
title Concordia: Self-Improving Synthetic Tables for Federated LLMs
topic Machine Learning
url https://arxiv.org/abs/2605.09855