Assessing Large Language Models on Climate Information

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Bulian, Jannis, Schäfer, Mike S., Amini, Afra, Lam, Heidi, Ciaramita, Massimiliano, Gaiarin, Ben, Hübscher, Michelle Chen, Buck, Christian, Mede, Niels G., Leippold, Markus, Strauß, Nadine
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917677260341248
author Bulian, Jannis
Schäfer, Mike S.
Amini, Afra
Lam, Heidi
Ciaramita, Massimiliano
Gaiarin, Ben
Hübscher, Michelle Chen
Buck, Christian
Mede, Niels G.
Leippold, Markus
Strauß, Nadine
author_facet Bulian, Jannis
Schäfer, Mike S.
Amini, Afra
Lam, Heidi
Ciaramita, Massimiliano
Gaiarin, Ben
Hübscher, Michelle Chen
Buck, Christian
Mede, Niels G.
Leippold, Markus
Strauß, Nadine
contents As Large Language Models (LLMs) rise in popularity, it is necessary to assess their capability in critically relevant domains. We present a comprehensive evaluation framework, grounded in science communication research, to assess LLM responses to questions about climate change. Our framework emphasizes both presentational and epistemological adequacy, offering a fine-grained analysis of LLM generations spanning 8 dimensions and 30 issues. Our evaluation task is a real-world example of a growing number of challenging problems where AI can complement and lift human performance. We introduce a novel protocol for scalable oversight that relies on AI Assistance and raters with relevant education. We evaluate several recent LLMs on a set of diverse climate questions. Our results point to a significant gap between surface and epistemological qualities of LLMs in the realm of climate communication.
format Preprint
id arxiv_https___arxiv_org_abs_2310_02932
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Assessing Large Language Models on Climate Information
Bulian, Jannis
Schäfer, Mike S.
Amini, Afra
Lam, Heidi
Ciaramita, Massimiliano
Gaiarin, Ben
Hübscher, Michelle Chen
Buck, Christian
Mede, Niels G.
Leippold, Markus
Strauß, Nadine
Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
As Large Language Models (LLMs) rise in popularity, it is necessary to assess their capability in critically relevant domains. We present a comprehensive evaluation framework, grounded in science communication research, to assess LLM responses to questions about climate change. Our framework emphasizes both presentational and epistemological adequacy, offering a fine-grained analysis of LLM generations spanning 8 dimensions and 30 issues. Our evaluation task is a real-world example of a growing number of challenging problems where AI can complement and lift human performance. We introduce a novel protocol for scalable oversight that relies on AI Assistance and raters with relevant education. We evaluate several recent LLMs on a set of diverse climate questions. Our results point to a significant gap between surface and epistemological qualities of LLMs in the realm of climate communication.
title Assessing Large Language Models on Climate Information
topic Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
url https://arxiv.org/abs/2310.02932