Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ramesh, Krithika, Smolyak, Daniel, Zhao, Zihao, Gandhi, Nupoor, Agarwal, Ritu, Bjarnadóttir, Margrét, Field, Anjalie
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2507.07229
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914129127669760
author Ramesh, Krithika
Smolyak, Daniel
Zhao, Zihao
Gandhi, Nupoor
Agarwal, Ritu
Bjarnadóttir, Margrét
Field, Anjalie
author_facet Ramesh, Krithika
Smolyak, Daniel
Zhao, Zihao
Gandhi, Nupoor
Agarwal, Ritu
Bjarnadóttir, Margrét
Field, Anjalie
contents We present SynthTextEval, a toolkit for conducting comprehensive evaluations of synthetic text. The fluency of large language model (LLM) outputs has made synthetic text potentially viable for numerous applications, such as reducing the risks of privacy violations in the development and deployment of AI systems in high-stakes domains. Realizing this potential, however, requires principled consistent evaluations of synthetic data across multiple dimensions: its utility in downstream systems, the fairness of these systems, the risk of privacy leakage, general distributional differences from the source text, and qualitative feedback from domain experts. SynthTextEval allows users to conduct evaluations along all of these dimensions over synthetic data that they upload or generate using the toolkit's generation module. While our toolkit can be run over any data, we highlight its functionality and effectiveness over datasets from two high-stakes domains: healthcare and law. By consolidating and standardizing evaluation metrics, we aim to improve the viability of synthetic text, and in-turn, privacy-preservation in AI development.
format Preprint
id arxiv_https___arxiv_org_abs_2507_07229
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SynthTextEval: Synthetic Text Data Generation and Evaluation for High-Stakes Domains
Ramesh, Krithika
Smolyak, Daniel
Zhao, Zihao
Gandhi, Nupoor
Agarwal, Ritu
Bjarnadóttir, Margrét
Field, Anjalie
Computation and Language
We present SynthTextEval, a toolkit for conducting comprehensive evaluations of synthetic text. The fluency of large language model (LLM) outputs has made synthetic text potentially viable for numerous applications, such as reducing the risks of privacy violations in the development and deployment of AI systems in high-stakes domains. Realizing this potential, however, requires principled consistent evaluations of synthetic data across multiple dimensions: its utility in downstream systems, the fairness of these systems, the risk of privacy leakage, general distributional differences from the source text, and qualitative feedback from domain experts. SynthTextEval allows users to conduct evaluations along all of these dimensions over synthetic data that they upload or generate using the toolkit's generation module. While our toolkit can be run over any data, we highlight its functionality and effectiveness over datasets from two high-stakes domains: healthcare and law. By consolidating and standardizing evaluation metrics, we aim to improve the viability of synthetic text, and in-turn, privacy-preservation in AI development.
title SynthTextEval: Synthetic Text Data Generation and Evaluation for High-Stakes Domains
topic Computation and Language
url https://arxiv.org/abs/2507.07229