Auto-ARGUE: LLM-Based Report Generation Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918473547907072 |
|---|---|
| author | Walden, William Mason, Marc Weller, Orion Dietz, Laura Conroy, John Molino, Neil Recknor, Hannah Li, Bryan Liu, Gabrielle Kaili-May Hou, Yu Lawrie, Dawn Mayfield, James Yang, Eugene |
| author_facet | Walden, William Mason, Marc Weller, Orion Dietz, Laura Conroy, John Molino, Neil Recknor, Hannah Li, Bryan Liu, Gabrielle Kaili-May Hou, Yu Lawrie, Dawn Mayfield, James Yang, Eugene |
| contents | Generation of citation-backed reports is a primary use case for retrieval-augmented generation (RAG) systems. While open-source evaluation tools exist for various RAG tasks, tools designed for report generation are lacking. Accordingly, we introduce Auto-ARGUE, a robust LLM-based implementation of the recently proposed ARGUE framework for report generation evaluation. We present analysis of Auto-ARGUE on the report generation pilot task from the TREC 2024 NeuCLIR track and on two tasks from the TREC 2024 RAG track, showing good system-level correlations with human judgments. Additionally, we release ARGUE-Viz, a web app for visualization and fine-grained analysis of Auto-ARGUE judgments and scores. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_26184 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Auto-ARGUE: LLM-Based Report Generation Evaluation Walden, William Mason, Marc Weller, Orion Dietz, Laura Conroy, John Molino, Neil Recknor, Hannah Li, Bryan Liu, Gabrielle Kaili-May Hou, Yu Lawrie, Dawn Mayfield, James Yang, Eugene Information Retrieval Artificial Intelligence Computation and Language Generation of citation-backed reports is a primary use case for retrieval-augmented generation (RAG) systems. While open-source evaluation tools exist for various RAG tasks, tools designed for report generation are lacking. Accordingly, we introduce Auto-ARGUE, a robust LLM-based implementation of the recently proposed ARGUE framework for report generation evaluation. We present analysis of Auto-ARGUE on the report generation pilot task from the TREC 2024 NeuCLIR track and on two tasks from the TREC 2024 RAG track, showing good system-level correlations with human judgments. Additionally, we release ARGUE-Viz, a web app for visualization and fine-grained analysis of Auto-ARGUE judgments and scores. |
| title | Auto-ARGUE: LLM-Based Report Generation Evaluation |
| topic | Information Retrieval Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2509.26184 |