Enhancing Long Document Long Form Summarisation with Self-Planning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866909970713280512 |
|---|---|
| author | Du, Xiaotang Saxena, Rohit Perez-Beltrachini, Laura Minervini, Pasquale Titov, Ivan |
| author_facet | Du, Xiaotang Saxena, Rohit Perez-Beltrachini, Laura Minervini, Pasquale Titov, Ivan |
| contents | We introduce a novel approach for long context summarisation, highlight-guided generation, that leverages sentence-level information as a content plan to improve the traceability and faithfulness of generated summaries. Our framework applies self-planning methods to identify important content and then generates a summary conditioned on the plan. We explore both an end-to-end and two-stage variants of the approach, finding that the two-stage pipeline performs better on long and information-dense documents. Experiments on long-form summarisation datasets demonstrate that our method consistently improves factual consistency while preserving relevance and overall quality. On GovReport, our best approach has improved ROUGE-L by 4.1 points and achieves about 35% gains in SummaC scores. Qualitative analysis shows that highlight-guided summarisation helps preserve important details, leading to more accurate and insightful summaries across domains. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_17179 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Enhancing Long Document Long Form Summarisation with Self-Planning Du, Xiaotang Saxena, Rohit Perez-Beltrachini, Laura Minervini, Pasquale Titov, Ivan Computation and Language We introduce a novel approach for long context summarisation, highlight-guided generation, that leverages sentence-level information as a content plan to improve the traceability and faithfulness of generated summaries. Our framework applies self-planning methods to identify important content and then generates a summary conditioned on the plan. We explore both an end-to-end and two-stage variants of the approach, finding that the two-stage pipeline performs better on long and information-dense documents. Experiments on long-form summarisation datasets demonstrate that our method consistently improves factual consistency while preserving relevance and overall quality. On GovReport, our best approach has improved ROUGE-L by 4.1 points and achieves about 35% gains in SummaC scores. Qualitative analysis shows that highlight-guided summarisation helps preserve important details, leading to more accurate and insightful summaries across domains. |
| title | Enhancing Long Document Long Form Summarisation with Self-Planning |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2512.17179 |