Saved in:
Bibliographic Details
Main Authors: Roemmele, Melissa, Gordon, Andrew S.
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2410.14897
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914979078209536
author Roemmele, Melissa
Gordon, Andrew S.
author_facet Roemmele, Melissa
Gordon, Andrew S.
contents LLMs can now perform a variety of complex writing tasks. They also excel in answering questions pertaining to natural language inference and commonsense reasoning. Composing these questions is itself a skilled writing task, so in this paper we consider LLMs as authors of commonsense assessment items. We prompt LLMs to generate items in the style of a prominent benchmark for commonsense reasoning, the Choice of Plausible Alternatives (COPA). We examine the outcome according to analyses facilitated by the LLMs and human annotation. We find that LLMs that succeed in answering the original COPA benchmark are also more successful in authoring their own items.
format Preprint
id arxiv_https___arxiv_org_abs_2410_14897
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle From Test-Taking to Test-Making: Examining LLM Authoring of Commonsense Assessment Items
Roemmele, Melissa
Gordon, Andrew S.
Computation and Language
Artificial Intelligence
LLMs can now perform a variety of complex writing tasks. They also excel in answering questions pertaining to natural language inference and commonsense reasoning. Composing these questions is itself a skilled writing task, so in this paper we consider LLMs as authors of commonsense assessment items. We prompt LLMs to generate items in the style of a prominent benchmark for commonsense reasoning, the Choice of Plausible Alternatives (COPA). We examine the outcome according to analyses facilitated by the LLMs and human annotation. We find that LLMs that succeed in answering the original COPA benchmark are also more successful in authoring their own items.
title From Test-Taking to Test-Making: Examining LLM Authoring of Commonsense Assessment Items
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2410.14897