Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2508.17973 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918132444037120 |
|---|---|
| author | Anschütz, Miriam Pham, Thanh Mai Nasrallah, Eslam Müller, Maximilian Craciun, Cristian-George Groh, Georg |
| author_facet | Anschütz, Miriam Pham, Thanh Mai Nasrallah, Eslam Müller, Maximilian Craciun, Cristian-George Groh, Georg |
| contents | The ability to paraphrase texts across different complexity levels is essential for creating accessible texts that can be tailored toward diverse reader groups. Thus, we introduce German4All, the first large-scale German dataset of aligned readability-controlled, paragraph-level paraphrases. It spans five readability levels and comprises over 25,000 samples. The dataset is automatically synthesized using GPT-4 and rigorously evaluated through both human and LLM-based judgments. Using German4All, we train an open-source, readability-controlled paraphrasing model that achieves state-of-the-art performance in German text simplification, enabling more nuanced and reader-specific adaptations. We opensource both the dataset and the model to encourage further research on multi-level paraphrasing |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_17973 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | German4All -- A Dataset and Model for Readability-Controlled Paraphrasing in German Anschütz, Miriam Pham, Thanh Mai Nasrallah, Eslam Müller, Maximilian Craciun, Cristian-George Groh, Georg Computation and Language The ability to paraphrase texts across different complexity levels is essential for creating accessible texts that can be tailored toward diverse reader groups. Thus, we introduce German4All, the first large-scale German dataset of aligned readability-controlled, paragraph-level paraphrases. It spans five readability levels and comprises over 25,000 samples. The dataset is automatically synthesized using GPT-4 and rigorously evaluated through both human and LLM-based judgments. Using German4All, we train an open-source, readability-controlled paraphrasing model that achieves state-of-the-art performance in German text simplification, enabling more nuanced and reader-specific adaptations. We opensource both the dataset and the model to encourage further research on multi-level paraphrasing |
| title | German4All -- A Dataset and Model for Readability-Controlled Paraphrasing in German |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2508.17973 |