Saved in:
Bibliographic Details
Main Authors: Anschütz, Miriam, Pham, Thanh Mai, Nasrallah, Eslam, Müller, Maximilian, Craciun, Cristian-George, Groh, Georg
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2508.17973
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918132444037120
author Anschütz, Miriam
Pham, Thanh Mai
Nasrallah, Eslam
Müller, Maximilian
Craciun, Cristian-George
Groh, Georg
author_facet Anschütz, Miriam
Pham, Thanh Mai
Nasrallah, Eslam
Müller, Maximilian
Craciun, Cristian-George
Groh, Georg
contents The ability to paraphrase texts across different complexity levels is essential for creating accessible texts that can be tailored toward diverse reader groups. Thus, we introduce German4All, the first large-scale German dataset of aligned readability-controlled, paragraph-level paraphrases. It spans five readability levels and comprises over 25,000 samples. The dataset is automatically synthesized using GPT-4 and rigorously evaluated through both human and LLM-based judgments. Using German4All, we train an open-source, readability-controlled paraphrasing model that achieves state-of-the-art performance in German text simplification, enabling more nuanced and reader-specific adaptations. We opensource both the dataset and the model to encourage further research on multi-level paraphrasing
format Preprint
id arxiv_https___arxiv_org_abs_2508_17973
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle German4All -- A Dataset and Model for Readability-Controlled Paraphrasing in German
Anschütz, Miriam
Pham, Thanh Mai
Nasrallah, Eslam
Müller, Maximilian
Craciun, Cristian-George
Groh, Georg
Computation and Language
The ability to paraphrase texts across different complexity levels is essential for creating accessible texts that can be tailored toward diverse reader groups. Thus, we introduce German4All, the first large-scale German dataset of aligned readability-controlled, paragraph-level paraphrases. It spans five readability levels and comprises over 25,000 samples. The dataset is automatically synthesized using GPT-4 and rigorously evaluated through both human and LLM-based judgments. Using German4All, we train an open-source, readability-controlled paraphrasing model that achieves state-of-the-art performance in German text simplification, enabling more nuanced and reader-specific adaptations. We opensource both the dataset and the model to encourage further research on multi-level paraphrasing
title German4All -- A Dataset and Model for Readability-Controlled Paraphrasing in German
topic Computation and Language
url https://arxiv.org/abs/2508.17973