Theoretical Physics Benchmark (TPBench) -- a Dataset and Study of AI Reasoning Capabilities in Theoretical Physics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chung, Daniel J. H., Gao, Zhiqi, Kvasiuk, Yurii, Li, Tianyi, Münchmeyer, Moritz, Rudolph, Maja, Sala, Frederic, Tadepalli, Sai Chaitanya
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910840478760960
author Chung, Daniel J. H.
Gao, Zhiqi
Kvasiuk, Yurii
Li, Tianyi
Münchmeyer, Moritz
Rudolph, Maja
Sala, Frederic
Tadepalli, Sai Chaitanya
author_facet Chung, Daniel J. H.
Gao, Zhiqi
Kvasiuk, Yurii
Li, Tianyi
Münchmeyer, Moritz
Rudolph, Maja
Sala, Frederic
Tadepalli, Sai Chaitanya
contents We introduce a benchmark to evaluate the capability of AI to solve problems in theoretical physics, focusing on high-energy theory and cosmology. The first iteration of our benchmark consists of 57 problems of varying difficulty, from undergraduate to research level. These problems are novel in the sense that they do not come from public problem collections. We evaluate our data set on various open and closed language models, including o3-mini, o1, DeepSeek-R1, GPT-4o and versions of Llama and Qwen. While we find impressive progress in model performance with the most recent models, our research-level difficulty problems are mostly unsolved. We address challenges of auto-verifiability and grading, and discuss common failure modes. While currently state-of-the art models are still of limited use for researchers, our results show that AI assisted theoretical physics research may become possible in the near future. We discuss the main obstacles towards this goal and possible strategies to overcome them. The public problems and solutions, results for various models, and updates to the data set and score distribution, are available on the website of the dataset tpbench.org.
format Preprint
id arxiv_https___arxiv_org_abs_2502_15815
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Theoretical Physics Benchmark (TPBench) -- a Dataset and Study of AI Reasoning Capabilities in Theoretical Physics
Chung, Daniel J. H.
Gao, Zhiqi
Kvasiuk, Yurii
Li, Tianyi
Münchmeyer, Moritz
Rudolph, Maja
Sala, Frederic
Tadepalli, Sai Chaitanya
Machine Learning
Cosmology and Nongalactic Astrophysics
Artificial Intelligence
High Energy Physics - Phenomenology
High Energy Physics - Theory
We introduce a benchmark to evaluate the capability of AI to solve problems in theoretical physics, focusing on high-energy theory and cosmology. The first iteration of our benchmark consists of 57 problems of varying difficulty, from undergraduate to research level. These problems are novel in the sense that they do not come from public problem collections. We evaluate our data set on various open and closed language models, including o3-mini, o1, DeepSeek-R1, GPT-4o and versions of Llama and Qwen. While we find impressive progress in model performance with the most recent models, our research-level difficulty problems are mostly unsolved. We address challenges of auto-verifiability and grading, and discuss common failure modes. While currently state-of-the art models are still of limited use for researchers, our results show that AI assisted theoretical physics research may become possible in the near future. We discuss the main obstacles towards this goal and possible strategies to overcome them. The public problems and solutions, results for various models, and updates to the data set and score distribution, are available on the website of the dataset tpbench.org.
title Theoretical Physics Benchmark (TPBench) -- a Dataset and Study of AI Reasoning Capabilities in Theoretical Physics
topic Machine Learning
Cosmology and Nongalactic Astrophysics
Artificial Intelligence
High Energy Physics - Phenomenology
High Energy Physics - Theory
url https://arxiv.org/abs/2502.15815