Saved in:
Bibliographic Details
Main Authors: Ezra, Elon, Weizman, Ariel, Azaria, Amos
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2508.12277
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908492692979712
author Ezra, Elon
Weizman, Ariel
Azaria, Amos
author_facet Ezra, Elon
Weizman, Ariel
Azaria, Amos
contents Large language models (LLMs) are commonly evaluated on tasks that test their knowledge or reasoning abilities. In this paper, we explore a different type of evaluation: whether an LLM can predict aspects of its own responses. Since LLMs lack the ability to execute themselves, we introduce the Self-Execution Benchmark, which measures a model's ability to anticipate properties of its output, such as whether a question will be difficult for it, whether it will refuse to answer, or what kinds of associations it is likely to produce. Our experiments show that models generally perform poorly on this benchmark, and that increased model size or capability does not consistently lead to better performance. These results suggest a fundamental limitation in how LLMs represent and reason about their own behavior.
format Preprint
id arxiv_https___arxiv_org_abs_2508_12277
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Self-Execution Benchmark: Measuring LLMs' Attempts to Overcome Their Lack of Self-Execution
Ezra, Elon
Weizman, Ariel
Azaria, Amos
Computation and Language
Artificial Intelligence
Large language models (LLMs) are commonly evaluated on tasks that test their knowledge or reasoning abilities. In this paper, we explore a different type of evaluation: whether an LLM can predict aspects of its own responses. Since LLMs lack the ability to execute themselves, we introduce the Self-Execution Benchmark, which measures a model's ability to anticipate properties of its output, such as whether a question will be difficult for it, whether it will refuse to answer, or what kinds of associations it is likely to produce. Our experiments show that models generally perform poorly on this benchmark, and that increased model size or capability does not consistently lead to better performance. These results suggest a fundamental limitation in how LLMs represent and reason about their own behavior.
title The Self-Execution Benchmark: Measuring LLMs' Attempts to Overcome Their Lack of Self-Execution
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2508.12277