FedSelfPlay-Defense: A Federated Adversarial Self-Play Framework for LLM Safety
Fuente:
Zenodo
Saved in:
| Main Author: | Jayathilaka, Hasini |
|---|---|
| Format: | Recurso digital |
| Published: |
Zenodo
2025
|
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Privacy-Preserving Prompt Injection Detection for LLMs Using Federated Learning and Embedding-Based NLP Classification
by: Jayathilaka, Hasini
Published: (2025)
by: Jayathilaka, Hasini
Published: (2025)
Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning
by: Li, Shangzhe, et al.
Published: (2026)
by: Li, Shangzhe, et al.
Published: (2026)
Seirênes: Adversarial Self-Play with Evolving Distractions for LLM Reasoning
by: Zhang, Chi, et al.
Published: (2026)
by: Zhang, Chi, et al.
Published: (2026)
TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment
by: Tan, Zhewen, et al.
Published: (2026)
by: Tan, Zhewen, et al.
Published: (2026)
Disentangling Intent from Role: Adversarial Self-Play for Persona-Invariant Safety Alignment
by: Li, Jiajia, et al.
Published: (2026)
by: Li, Jiajia, et al.
Published: (2026)
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
by: Chen, Jiaqi, et al.
Published: (2025)
by: Chen, Jiaqi, et al.
Published: (2025)
Learning Robust Reasoning through Guided Adversarial Self-Play
by: Li, Shuozhe, et al.
Published: (2026)
by: Li, Shuozhe, et al.
Published: (2026)
ASP2LJ : An Adversarial Self-Play Laywer Augmented Legal Judgment Framework
by: Chang, Ao, et al.
Published: (2025)
by: Chang, Ao, et al.
Published: (2025)
Model Behavior Specification by Leveraging LLM Self-Playing and Self-Improving
by: Park, Soya, et al.
Published: (2025)
by: Park, Soya, et al.
Published: (2025)
Scaling Self-Play with Self-Guidance
by: Bailey, Luke, et al.
Published: (2026)
by: Bailey, Luke, et al.
Published: (2026)
Stay in Character, Stay Safe: Dual-Cycle Adversarial Self-Evolution for Safety Role-Playing Agents
by: Liao, Mingyang, et al.
Published: (2026)
by: Liao, Mingyang, et al.
Published: (2026)
SPRec: Self-Play to Debias LLM-based Recommendation
by: Gao, Chongming, et al.
Published: (2024)
by: Gao, Chongming, et al.
Published: (2024)
The Attacker in the Mirror: Breaking Self-Consistency in Safety via Anchored Bipolicy Self-Play
by: La Malfa, Gabriele, et al.
Published: (2026)
by: La Malfa, Gabriele, et al.
Published: (2026)
SPARS: Self-Play Adversarial Reinforcement Learning for Segmentation of Liver Tumours
by: Tan, Catalina, et al.
Published: (2025)
by: Tan, Catalina, et al.
Published: (2025)
A Theoretical Framework for Self-Play Theorem Proving Algorithms
by: Chen, Thomas, et al.
Published: (2026)
by: Chen, Thomas, et al.
Published: (2026)
Enhancing Language Agent Strategic Reasoning through Self-Play in Adversarial Games
by: Zhang, Yikai, et al.
Published: (2025)
by: Zhang, Yikai, et al.
Published: (2025)
When Actions Disappear: Adversarial Action Removal in Self-Play Reinforcement Learning
by: Kujur, Arahan
Published: (2026)
by: Kujur, Arahan
Published: (2026)
Self-Play with Adversarial Critic: Provable and Scalable Offline Alignment for Language Models
by: Ji, Xiang, et al.
Published: (2024)
by: Ji, Xiang, et al.
Published: (2024)
$π$-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data
by: Zhang, Yaocheng, et al.
Published: (2026)
by: Zhang, Yaocheng, et al.
Published: (2026)
Self-Improving AI Agents through Self-Play
by: Chojecki, Przemyslaw
Published: (2025)
by: Chojecki, Przemyslaw
Published: (2025)
R-Diverse: Mitigating Diversity Illusion in Self-Play LLM Training
by: Li, Gengsheng, et al.
Published: (2026)
by: Li, Gengsheng, et al.
Published: (2026)
Learning to Solve and Verify: A Self-Play Framework for Code and Test Generation
by: Lin, Zi, et al.
Published: (2025)
by: Lin, Zi, et al.
Published: (2025)
Can Large Language Models Play Games? A Case Study of A Self-Play Approach
by: Guo, Hongyi, et al.
Published: (2024)
by: Guo, Hongyi, et al.
Published: (2024)
Language Self-Play For Data-Free Training
by: Kuba, Jakub Grudzien, et al.
Published: (2025)
by: Kuba, Jakub Grudzien, et al.
Published: (2025)
Robust Autonomy Emerges from Self-Play
by: Cusumano-Towner, Marco, et al.
Published: (2025)
by: Cusumano-Towner, Marco, et al.
Published: (2025)
Meta-Learning in Self-Play Regret Minimization
by: Sychrovský, David, et al.
Published: (2025)
by: Sychrovský, David, et al.
Published: (2025)
Investigating Regularization of Self-Play Language Models
by: Alami, Reda, et al.
Published: (2024)
by: Alami, Reda, et al.
Published: (2024)
Offline Fictitious Self-Play for Competitive Games
by: Chen, Jingxiao, et al.
Published: (2024)
by: Chen, Jingxiao, et al.
Published: (2024)
Differentially Private Reinforcement Learning with Self-Play
by: Qiao, Dan, et al.
Published: (2024)
by: Qiao, Dan, et al.
Published: (2024)
Learning to Drive via Asymmetric Self-Play
by: Zhang, Chris, et al.
Published: (2024)
by: Zhang, Chris, et al.
Published: (2024)
PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play
by: Castanyer, Roger Creus, et al.
Published: (2026)
by: Castanyer, Roger Creus, et al.
Published: (2026)
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
by: Spiliopoulou, Evangelia, et al.
Published: (2025)
by: Spiliopoulou, Evangelia, et al.
Published: (2025)
FedBAP: Backdoor Defense via Benign Adversarial Perturbation in Federated Learning
by: Yan, Xinhai, et al.
Published: (2025)
by: Yan, Xinhai, et al.
Published: (2025)
Self-Play Enhancement via Advantage-Weighted Refinement in Online Federated LLM Fine-Tuning with Real-Time Feedback
by: Lee, Seohyun, et al.
Published: (2026)
by: Lee, Seohyun, et al.
Published: (2026)
Be Your Own Red Teamer: Safety Alignment via Self-Play and Reflective Experience Replay
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
Improving LLM Code Reasoning via Semantic Equivalence Self-Play with Formal Verification
by: Barone, Antonio Valerio Miceli, et al.
Published: (2026)
by: Barone, Antonio Valerio Miceli, et al.
Published: (2026)
A Tale of Too Many Doctrines: Supervening Impossibility and the Sale of Goods
by: Chathuni Jayathilaka
Published: (2024)
by: Chathuni Jayathilaka
Published: (2024)
Heterogeneous Adversarial Play in Interactive Environments
by: Xu, Manjie, et al.
Published: (2025)
by: Xu, Manjie, et al.
Published: (2025)
Conversational Self-Play for Discovering and Understanding Psychotherapy Approaches
by: Kampman, Onno P, et al.
Published: (2025)
by: Kampman, Onno P, et al.
Published: (2025)
SPICE: Self-Play In Corpus Environments Improves Reasoning
by: Liu, Bo, et al.
Published: (2025)
by: Liu, Bo, et al.
Published: (2025)
Similar Items
-
Privacy-Preserving Prompt Injection Detection for LLMs Using Federated Learning and Embedding-Based NLP Classification
by: Jayathilaka, Hasini
Published: (2025) -
Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning
by: Li, Shangzhe, et al.
Published: (2026) -
Seirênes: Adversarial Self-Play with Evolving Distractions for LLM Reasoning
by: Zhang, Chi, et al.
Published: (2026) -
TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment
by: Tan, Zhewen, et al.
Published: (2026) -
Disentangling Intent from Role: Adversarial Self-Play for Persona-Invariant Safety Alignment
by: Li, Jiajia, et al.
Published: (2026)