Assessing Large Language Models on Climate Information
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866917677260341248 |
|---|---|
| author | Bulian, Jannis Schäfer, Mike S. Amini, Afra Lam, Heidi Ciaramita, Massimiliano Gaiarin, Ben Hübscher, Michelle Chen Buck, Christian Mede, Niels G. Leippold, Markus Strauß, Nadine |
| author_facet | Bulian, Jannis Schäfer, Mike S. Amini, Afra Lam, Heidi Ciaramita, Massimiliano Gaiarin, Ben Hübscher, Michelle Chen Buck, Christian Mede, Niels G. Leippold, Markus Strauß, Nadine |
| contents | As Large Language Models (LLMs) rise in popularity, it is necessary to assess their capability in critically relevant domains. We present a comprehensive evaluation framework, grounded in science communication research, to assess LLM responses to questions about climate change. Our framework emphasizes both presentational and epistemological adequacy, offering a fine-grained analysis of LLM generations spanning 8 dimensions and 30 issues. Our evaluation task is a real-world example of a growing number of challenging problems where AI can complement and lift human performance. We introduce a novel protocol for scalable oversight that relies on AI Assistance and raters with relevant education. We evaluate several recent LLMs on a set of diverse climate questions. Our results point to a significant gap between surface and epistemological qualities of LLMs in the realm of climate communication. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2310_02932 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Assessing Large Language Models on Climate Information Bulian, Jannis Schäfer, Mike S. Amini, Afra Lam, Heidi Ciaramita, Massimiliano Gaiarin, Ben Hübscher, Michelle Chen Buck, Christian Mede, Niels G. Leippold, Markus Strauß, Nadine Computation and Language Artificial Intelligence Computers and Society Machine Learning As Large Language Models (LLMs) rise in popularity, it is necessary to assess their capability in critically relevant domains. We present a comprehensive evaluation framework, grounded in science communication research, to assess LLM responses to questions about climate change. Our framework emphasizes both presentational and epistemological adequacy, offering a fine-grained analysis of LLM generations spanning 8 dimensions and 30 issues. Our evaluation task is a real-world example of a growing number of challenging problems where AI can complement and lift human performance. We introduce a novel protocol for scalable oversight that relies on AI Assistance and raters with relevant education. We evaluate several recent LLMs on a set of diverse climate questions. Our results point to a significant gap between surface and epistemological qualities of LLMs in the realm of climate communication. |
| title | Assessing Large Language Models on Climate Information |
| topic | Computation and Language Artificial Intelligence Computers and Society Machine Learning |
| url | https://arxiv.org/abs/2310.02932 |