Responsible Federated LLMs via Safety Filtering and Constitutional AI

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Noh, Eunchung, Baek, Jeonghun
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918506841243648
author Noh, Eunchung
Baek, Jeonghun
author_facet Noh, Eunchung
Baek, Jeonghun
contents Recent research has increasingly focused on training large language models (LLMs) using federated learning, known as FedLLM. However, responsible AI (RAI), which aims to ensure safe and trustworthy responses, remains underexplored in this context. In FedLLM, client-side training data may contain harmful content, resulting in unsafe LLMs that can generate inappropriate responses. Aggregating such models into a global model and redistributing it to clients risks the widespread deployment of unsafe LLMs. To address this, we incorporate two well-established RAI techniques into FedLLM: safety filtering and constitutional AI. Our experiments show that these methods significantly improve LLM safety, achieving over 20% improvement on AdvBench.
format Preprint
id arxiv_https___arxiv_org_abs_2502_16691
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Responsible Federated LLMs via Safety Filtering and Constitutional AI
Noh, Eunchung
Baek, Jeonghun
Computation and Language
Distributed, Parallel, and Cluster Computing
Multiagent Systems
Recent research has increasingly focused on training large language models (LLMs) using federated learning, known as FedLLM. However, responsible AI (RAI), which aims to ensure safe and trustworthy responses, remains underexplored in this context. In FedLLM, client-side training data may contain harmful content, resulting in unsafe LLMs that can generate inappropriate responses. Aggregating such models into a global model and redistributing it to clients risks the widespread deployment of unsafe LLMs. To address this, we incorporate two well-established RAI techniques into FedLLM: safety filtering and constitutional AI. Our experiments show that these methods significantly improve LLM safety, achieving over 20% improvement on AdvBench.
title Responsible Federated LLMs via Safety Filtering and Constitutional AI
topic Computation and Language
Distributed, Parallel, and Cluster Computing
Multiagent Systems
url https://arxiv.org/abs/2502.16691