Towards a Benchmark for Scientific Understanding in Humans and Machines

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Barman, Kristian Gonzalez, Caron, Sascha, Claassen, Tom, de Regt, Henk
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914784767639552
author Barman, Kristian Gonzalez
Caron, Sascha
Claassen, Tom
de Regt, Henk
author_facet Barman, Kristian Gonzalez
Caron, Sascha
Claassen, Tom
de Regt, Henk
contents Scientific understanding is a fundamental goal of science, allowing us to explain the world. There is currently no good way to measure the scientific understanding of agents, whether these be humans or Artificial Intelligence systems. Without a clear benchmark, it is challenging to evaluate and compare different levels of and approaches to scientific understanding. In this Roadmap, we propose a framework to create a benchmark for scientific understanding, utilizing tools from philosophy of science. We adopt a behavioral notion according to which genuine understanding should be recognized as an ability to perform certain tasks. We extend this notion by considering a set of questions that can gauge different levels of scientific understanding, covering information retrieval, the capability to arrange information to produce an explanation, and the ability to infer how things would be different under different circumstances. The Scientific Understanding Benchmark (SUB), which is formed by a set of these tests, allows for the evaluation and comparison of different approaches. Benchmarking plays a crucial role in establishing trust, ensuring quality control, and providing a basis for performance evaluation. By aligning machine and human scientific understanding we can improve their utility, ultimately advancing scientific understanding and helping to discover new insights within machines.
format Preprint
id arxiv_https___arxiv_org_abs_2304_10327
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Towards a Benchmark for Scientific Understanding in Humans and Machines
Barman, Kristian Gonzalez
Caron, Sascha
Claassen, Tom
de Regt, Henk
Artificial Intelligence
Computation and Language
Human-Computer Interaction
High Energy Physics - Phenomenology
History and Philosophy of Physics
Scientific understanding is a fundamental goal of science, allowing us to explain the world. There is currently no good way to measure the scientific understanding of agents, whether these be humans or Artificial Intelligence systems. Without a clear benchmark, it is challenging to evaluate and compare different levels of and approaches to scientific understanding. In this Roadmap, we propose a framework to create a benchmark for scientific understanding, utilizing tools from philosophy of science. We adopt a behavioral notion according to which genuine understanding should be recognized as an ability to perform certain tasks. We extend this notion by considering a set of questions that can gauge different levels of scientific understanding, covering information retrieval, the capability to arrange information to produce an explanation, and the ability to infer how things would be different under different circumstances. The Scientific Understanding Benchmark (SUB), which is formed by a set of these tests, allows for the evaluation and comparison of different approaches. Benchmarking plays a crucial role in establishing trust, ensuring quality control, and providing a basis for performance evaluation. By aligning machine and human scientific understanding we can improve their utility, ultimately advancing scientific understanding and helping to discover new insights within machines.
title Towards a Benchmark for Scientific Understanding in Humans and Machines
topic Artificial Intelligence
Computation and Language
Human-Computer Interaction
High Energy Physics - Phenomenology
History and Philosophy of Physics
url https://arxiv.org/abs/2304.10327