De-jargonizing Science for Journalists with GPT-4: A Pilot Study
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866929544930263040 |
|---|---|
| author | Nishal, Sachita Lee, Eric Diakopoulos, Nicholas |
| author_facet | Nishal, Sachita Lee, Eric Diakopoulos, Nicholas |
| contents | This study offers an initial evaluation of a human-in-the-loop system leveraging GPT-4 (a large language model or LLM), and Retrieval-Augmented Generation (RAG) to identify and define jargon terms in scientific abstracts, based on readers' self-reported knowledge. The system achieves fairly high recall in identifying jargon and preserves relative differences in readers' jargon identification, suggesting personalization as a feasible use-case for LLMs to support sense-making of complex information. Surprisingly, using only abstracts for context to generate definitions yields slightly more accurate and higher quality definitions than using RAG-based context from the fulltext of an article. The findings highlight the potential of generative AI for assisting science reporters, and can inform future work on developing tools to simplify dense documents. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_12069 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | De-jargonizing Science for Journalists with GPT-4: A Pilot Study Nishal, Sachita Lee, Eric Diakopoulos, Nicholas Computation and Language Computers and Society Human-Computer Interaction H.4; H.5 This study offers an initial evaluation of a human-in-the-loop system leveraging GPT-4 (a large language model or LLM), and Retrieval-Augmented Generation (RAG) to identify and define jargon terms in scientific abstracts, based on readers' self-reported knowledge. The system achieves fairly high recall in identifying jargon and preserves relative differences in readers' jargon identification, suggesting personalization as a feasible use-case for LLMs to support sense-making of complex information. Surprisingly, using only abstracts for context to generate definitions yields slightly more accurate and higher quality definitions than using RAG-based context from the fulltext of an article. The findings highlight the potential of generative AI for assisting science reporters, and can inform future work on developing tools to simplify dense documents. |
| title | De-jargonizing Science for Journalists with GPT-4: A Pilot Study |
| topic | Computation and Language Computers and Society Human-Computer Interaction H.4; H.5 |
| url | https://arxiv.org/abs/2410.12069 |