SpreadsheetArena: Decomposing Preference in LLM Generation of Spreadsheet Workbooks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kundurthy, Srivatsa, Na, Clara, Handley, Michael, Kirshner, Zach, Zhang, Chen Bo Calvin, Sharma, Manasi, Strubell, Emma, Ling, John
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915851562647552
author Kundurthy, Srivatsa
Na, Clara
Handley, Michael
Kirshner, Zach
Zhang, Chen Bo Calvin
Sharma, Manasi
Strubell, Emma
Ling, John
author_facet Kundurthy, Srivatsa
Na, Clara
Handley, Michael
Kirshner, Zach
Zhang, Chen Bo Calvin
Sharma, Manasi
Strubell, Emma
Ling, John
contents Large language models (LLMs) are increasingly tasked with producing and manipulating structured artifacts. We consider the task of end-to-end spreadsheet generation, where language models are prompted to produce spreadsheet artifacts to satisfy users' explicit and implicit constraints, specified in natural language. We introduce SpreadsheetArena, a platform for evaluating models' performance on the task via blind pairwise evaluations of LLM-generated spreadsheet workbooks. As with other complex, open-ended tasks, relevant evaluation criteria can vary substantially across use cases and prompts, often in ways that are difficult to formalize. Compared to general chat or text generation settings, spreadsheet generation presents unique challenges and opportunities: the task output structure is well-defined and multi-dimensional, and there are often complex considerations around interactivity and layout. Among other findings, we observe that stylistic, structural, and functional features of preferred spreadsheets vary substantially across use cases, and expert evaluations of spreadsheets for finance prompts suggests that even highly ranked arena models do not reliably produce spreadsheets aligned with domain-specific best practices. Our hope is that our work prompts further study of end-to-end spreadsheet generation as a challenging and interesting category of complex, open-ended tasks for LLMs. Our live arena is hosted at https://spreadsheetarena.ai.
format Preprint
id arxiv_https___arxiv_org_abs_2603_10002
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SpreadsheetArena: Decomposing Preference in LLM Generation of Spreadsheet Workbooks
Kundurthy, Srivatsa
Na, Clara
Handley, Michael
Kirshner, Zach
Zhang, Chen Bo Calvin
Sharma, Manasi
Strubell, Emma
Ling, John
Computation and Language
Artificial Intelligence
Machine Learning
Large language models (LLMs) are increasingly tasked with producing and manipulating structured artifacts. We consider the task of end-to-end spreadsheet generation, where language models are prompted to produce spreadsheet artifacts to satisfy users' explicit and implicit constraints, specified in natural language. We introduce SpreadsheetArena, a platform for evaluating models' performance on the task via blind pairwise evaluations of LLM-generated spreadsheet workbooks. As with other complex, open-ended tasks, relevant evaluation criteria can vary substantially across use cases and prompts, often in ways that are difficult to formalize. Compared to general chat or text generation settings, spreadsheet generation presents unique challenges and opportunities: the task output structure is well-defined and multi-dimensional, and there are often complex considerations around interactivity and layout. Among other findings, we observe that stylistic, structural, and functional features of preferred spreadsheets vary substantially across use cases, and expert evaluations of spreadsheets for finance prompts suggests that even highly ranked arena models do not reliably produce spreadsheets aligned with domain-specific best practices. Our hope is that our work prompts further study of end-to-end spreadsheet generation as a challenging and interesting category of complex, open-ended tasks for LLMs. Our live arena is hosted at https://spreadsheetarena.ai.
title SpreadsheetArena: Decomposing Preference in LLM Generation of Spreadsheet Workbooks
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2603.10002