Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Nagl, Sebastian, Elganayni, Mohamed, Pospisil, Melanie, Grabmair, Matthias
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2512.09434
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915665877663744
author Nagl, Sebastian
Elganayni, Mohamed
Pospisil, Melanie
Grabmair, Matthias
author_facet Nagl, Sebastian
Elganayni, Mohamed
Pospisil, Melanie
Grabmair, Matthias
contents Official court press releases from Germany's highest courts present and explain judicial rulings to the public, as well as to expert audiences. Prior NLP efforts emphasize technical headnotes, ignoring citizen-oriented communication needs. We introduce CourtPressGER, a 6.4k dataset of triples: rulings, human-drafted press releases, and synthetic prompts for LLMs to generate comparable releases. This benchmark trains and evaluates LLMs in generating accurate, readable summaries from long judicial texts. We benchmark small and large LLMs using reference-based metrics, factual-consistency checks, LLM-as-judge, and expert ranking. Large LLMs produce high-quality drafts with minimal hierarchical performance loss; smaller models require hierarchical setups for long judgments. Initial benchmarks show varying model performance, with human-drafted releases ranking highest.
format Preprint
id arxiv_https___arxiv_org_abs_2512_09434
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CourtPressGER: A German Court Decision to Press Release Summarization Dataset
Nagl, Sebastian
Elganayni, Mohamed
Pospisil, Melanie
Grabmair, Matthias
Computation and Language
Artificial Intelligence
Official court press releases from Germany's highest courts present and explain judicial rulings to the public, as well as to expert audiences. Prior NLP efforts emphasize technical headnotes, ignoring citizen-oriented communication needs. We introduce CourtPressGER, a 6.4k dataset of triples: rulings, human-drafted press releases, and synthetic prompts for LLMs to generate comparable releases. This benchmark trains and evaluates LLMs in generating accurate, readable summaries from long judicial texts. We benchmark small and large LLMs using reference-based metrics, factual-consistency checks, LLM-as-judge, and expert ranking. Large LLMs produce high-quality drafts with minimal hierarchical performance loss; smaller models require hierarchical setups for long judgments. Initial benchmarks show varying model performance, with human-drafted releases ranking highest.
title CourtPressGER: A German Court Decision to Press Release Summarization Dataset
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2512.09434