| _version_ | 1866901386542710784 |
|---|---|
| author | Jayathilaka, Hasini |
| author_facet | Jayathilaka, Hasini |
| contents | <p>Ensuring the safety and robustness of large language models (LLMs) remains a critical challenge, particularly in adversarial settings. This work introduces FedSelfPlay-Defense, a fully synthetic federated adversarial self-play framework for improving LLM safety. In this system, each client hosts two lightweight synthetic agents: an attacker agent that generates adversarial prompts through simple mutation-based evolution, and a defender agent that learns to classify and block harmful content. Because all prompts are synthetic and generated locally, no real user data is required. Only defender model updates and <em><span>collected (not aggregated)</span></em> attack embeddings are shared with the server, preserving privacy while enabling collaborative training. Experimental results show that the defender rapidly converges to a stable operating point, achieving high recall and maintaining a conservative bias typical of safety-oriented filters. t-SNE visualizations of collected attacker embeddings illustrate the diversity of locally generated attacks across clients. Although overall accuracy remains stable rather than progressively improving, the framework effectively demonstrates how federated self-play can support autonomous, privacy-preserving safety training. FedSelfPlay-Defense provides a scalable foundation for future work on stronger attackers, more expressive defenders, and richer forms of federated safety adaptation.</p> <p>e foundation for future work on stronger attackers, more expressive defenders, and richer forms of federated safety adaptation.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_17924331 |
| institution | Zenodo |
| language | |
| publishDate | 2025 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | FedSelfPlay-Defense: A Federated Adversarial Self-Play Framework for LLM Safety Jayathilaka, Hasini <p>Ensuring the safety and robustness of large language models (LLMs) remains a critical challenge, particularly in adversarial settings. This work introduces FedSelfPlay-Defense, a fully synthetic federated adversarial self-play framework for improving LLM safety. In this system, each client hosts two lightweight synthetic agents: an attacker agent that generates adversarial prompts through simple mutation-based evolution, and a defender agent that learns to classify and block harmful content. Because all prompts are synthetic and generated locally, no real user data is required. Only defender model updates and <em><span>collected (not aggregated)</span></em> attack embeddings are shared with the server, preserving privacy while enabling collaborative training. Experimental results show that the defender rapidly converges to a stable operating point, achieving high recall and maintaining a conservative bias typical of safety-oriented filters. t-SNE visualizations of collected attacker embeddings illustrate the diversity of locally generated attacks across clients. Although overall accuracy remains stable rather than progressively improving, the framework effectively demonstrates how federated self-play can support autonomous, privacy-preserving safety training. FedSelfPlay-Defense provides a scalable foundation for future work on stronger attackers, more expressive defenders, and richer forms of federated safety adaptation.</p> <p>e foundation for future work on stronger attackers, more expressive defenders, and richer forms of federated safety adaptation.</p> |
| title | FedSelfPlay-Defense: A Federated Adversarial Self-Play Framework for LLM Safety |
| url | https://doi.org/10.5281/zenodo.17924331 |