Adversarial Arena: Crowdsourcing Data Generation through Interactive Competition

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Goyal, Prasoon, Sahai, Sattvik, Johnston, Michael, Shi, Hangjie, Lu, Yao, Liu, Shaohua, Rumshisky, Anna, Gupta, Rahul, Gottardi, Anna, Zhang, Desheng, Vaz, Lavina, Ball, Leslie, Hu, Lucy, Dai, Luke, Sagi, Samyuth, Murray, Maureen, Ananthakrishnan, Sankaranarayanan
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908978612535296
author Goyal, Prasoon
Sahai, Sattvik
Johnston, Michael
Shi, Hangjie
Lu, Yao
Liu, Shaohua
Rumshisky, Anna
Gupta, Rahul
Gottardi, Anna
Zhang, Desheng
Vaz, Lavina
Ball, Leslie
Hu, Lucy
Dai, Luke
Sagi, Samyuth
Murray, Maureen
Ananthakrishnan, Sankaranarayanan
author_facet Goyal, Prasoon
Sahai, Sattvik
Johnston, Michael
Shi, Hangjie
Lu, Yao
Liu, Shaohua
Rumshisky, Anna
Gupta, Rahul
Gottardi, Anna
Zhang, Desheng
Vaz, Lavina
Ball, Leslie
Hu, Lucy
Dai, Luke
Sagi, Samyuth
Murray, Maureen
Ananthakrishnan, Sankaranarayanan
contents Post-training Large Language Models requires diverse, high-quality data which is rare and costly to obtain, especially in low resource domains and for multi-turn conversations. Common solutions are crowdsourcing or synthetic generation, but both often yield low-quality or low-diversity data. We introduce Adversarial Arena for building high quality conversational datasets by framing data generation as an adversarial task: attackers create prompts, and defenders generate responses. This interactive competition between multiple teams naturally produces diverse and complex data. We validated this approach by conducting a competition with 10 academic teams from top US and European universities, each building attacker or defender bots. The competition, focused on safety alignment of LLMs in cybersecurity, generated 19,683 multi-turn conversations. Fine-tuning an open-source model on this dataset produced an 18.47% improvement in secure code generation on CyberSecEval-Instruct and 29.42% improvement on CyberSecEval-MITRE.
format Preprint
id arxiv_https___arxiv_org_abs_2604_17803
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Adversarial Arena: Crowdsourcing Data Generation through Interactive Competition
Goyal, Prasoon
Sahai, Sattvik
Johnston, Michael
Shi, Hangjie
Lu, Yao
Liu, Shaohua
Rumshisky, Anna
Gupta, Rahul
Gottardi, Anna
Zhang, Desheng
Vaz, Lavina
Ball, Leslie
Hu, Lucy
Dai, Luke
Sagi, Samyuth
Murray, Maureen
Ananthakrishnan, Sankaranarayanan
Artificial Intelligence
Machine Learning
I.2.7; I.2.6; E.0
Post-training Large Language Models requires diverse, high-quality data which is rare and costly to obtain, especially in low resource domains and for multi-turn conversations. Common solutions are crowdsourcing or synthetic generation, but both often yield low-quality or low-diversity data. We introduce Adversarial Arena for building high quality conversational datasets by framing data generation as an adversarial task: attackers create prompts, and defenders generate responses. This interactive competition between multiple teams naturally produces diverse and complex data. We validated this approach by conducting a competition with 10 academic teams from top US and European universities, each building attacker or defender bots. The competition, focused on safety alignment of LLMs in cybersecurity, generated 19,683 multi-turn conversations. Fine-tuning an open-source model on this dataset produced an 18.47% improvement in secure code generation on CyberSecEval-Instruct and 29.42% improvement on CyberSecEval-MITRE.
title Adversarial Arena: Crowdsourcing Data Generation through Interactive Competition
topic Artificial Intelligence
Machine Learning
I.2.7; I.2.6; E.0
url https://arxiv.org/abs/2604.17803