NodeSynth: Socially Aligned Synthetic Data for AI Evaluation
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909051710865408 |
|---|---|
| author | Rashid, Qazi Mamunur Yang, Xuan Yang, Zhengzhe Pan, Yanzhou van Liemt, Erin Neal, Darlene Pancholi, Kshitij Smith-Loud, Jamila |
| author_facet | Rashid, Qazi Mamunur Yang, Xuan Yang, Zhengzhe Pan, Yanzhou van Liemt, Erin Neal, Darlene Pancholi, Kshitij Smith-Loud, Jamila |
| contents | Recent advancements in generative AI facilitate large-scale synthetic data generation for model evaluation. However, without targeted approaches, these datasets often lack the sociotechnical nuance required for sensitive domains. We introduce NodeSynth, an evidence-grounded methodology that generates socially relevant synthetic queries by leveraging a fine-tuned taxonomy generator (TaG) anchored in real-world evidence. Evaluated against four mainstream LLMs (e.g., Claude 4.5 Haiku), NodeSynth elicited failure rates up to five times higher than human-authored benchmarks. Ablation studies confirm that our granular taxonomic expansion significantly drives these failure rates, while independent validation reveals critical deficiencies in prominent guard models (e.g., Llama-Guard-3). We open-source our end-to-end research prototype and datasets to enable scalable, high-stakes model evaluation and targeted safety interventions (https://github.com/google-research/nodesynth). |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_14381 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | NodeSynth: Socially Aligned Synthetic Data for AI Evaluation Rashid, Qazi Mamunur Yang, Xuan Yang, Zhengzhe Pan, Yanzhou van Liemt, Erin Neal, Darlene Pancholi, Kshitij Smith-Loud, Jamila Machine Learning Computation and Language Recent advancements in generative AI facilitate large-scale synthetic data generation for model evaluation. However, without targeted approaches, these datasets often lack the sociotechnical nuance required for sensitive domains. We introduce NodeSynth, an evidence-grounded methodology that generates socially relevant synthetic queries by leveraging a fine-tuned taxonomy generator (TaG) anchored in real-world evidence. Evaluated against four mainstream LLMs (e.g., Claude 4.5 Haiku), NodeSynth elicited failure rates up to five times higher than human-authored benchmarks. Ablation studies confirm that our granular taxonomic expansion significantly drives these failure rates, while independent validation reveals critical deficiencies in prominent guard models (e.g., Llama-Guard-3). We open-source our end-to-end research prototype and datasets to enable scalable, high-stakes model evaluation and targeted safety interventions (https://github.com/google-research/nodesynth). |
| title | NodeSynth: Socially Aligned Synthetic Data for AI Evaluation |
| topic | Machine Learning Computation and Language |
| url | https://arxiv.org/abs/2605.14381 |