SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Zeyao, Zhang, Bohan, Zhang, Jing, Yu, Jifan, Zhang, Xiaokang, Zhang, Xiaohan, Luo, Sijia, Wang, Xi, Tang, Jie |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dynamic Scaling of Unit Tests for Code Reward Modeling
by: Ma, Zeyao, et al.
Published: (2025)
by: Ma, Zeyao, et al.
Published: (2025)
Classification of Spreadsheet Errors
by: Rajalingham, Kamalasen, et al.
Published: (2008)
by: Rajalingham, Kamalasen, et al.
Published: (2008)
Drivers of the Cost of Spreadsheet Audit
by: Colver, David
Published: (2011)
by: Colver, David
Published: (2011)
Consensus-Free Spreadsheet Integration
by: Baylor, Brandon, et al.
Published: (2022)
by: Baylor, Brandon, et al.
Published: (2022)
SODBench: A Large Language Model Approach to Documenting Spreadsheet Operations
by: Indika, Amila, et al.
Published: (2025)
by: Indika, Amila, et al.
Published: (2025)
User Defined Spreadsheet Functions in Excel
by: Tyszkiewicz, Jerzy, et al.
Published: (2012)
by: Tyszkiewicz, Jerzy, et al.
Published: (2012)
In Pursuit of Spreadsheet Excellence
by: Croll, Grenville J.
Published: (2008)
by: Croll, Grenville J.
Published: (2008)
Spreadsheet Debugging
by: Ayalew, Yirsaw, et al.
Published: (2008)
by: Ayalew, Yirsaw, et al.
Published: (2008)
Large Language Models for Spreadsheets: Benchmarking Progress and Evaluating Performance with FLARE
by: Thorne, Simon
Published: (2025)
by: Thorne, Simon
Published: (2025)
ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
by: Yang, Jie, et al.
Published: (2026)
by: Yang, Jie, et al.
Published: (2026)
Computational Models of Spreadsheet Development: Basis for Educational Approaches
by: Hodnigg, Karin, et al.
Published: (2008)
by: Hodnigg, Karin, et al.
Published: (2008)
Benchmarking LLMs for Unit Test Generation from Real-World Functions
by: Huang, Dong, et al.
Published: (2025)
by: Huang, Dong, et al.
Published: (2025)
NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts
by: Zhang, Shudan, et al.
Published: (2024)
by: Zhang, Shudan, et al.
Published: (2024)
CoT-based Synthesizer: Enhancing LLM Performance through Answer Synthesis
by: Zhang, Bohan, et al.
Published: (2025)
by: Zhang, Bohan, et al.
Published: (2025)
EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories
by: Li, Jia, et al.
Published: (2024)
by: Li, Jia, et al.
Published: (2024)
Spreadsheet Modeling Experiments Using GPTs on Small Problem Statements and the Wall Task
by: Grossman, Thomas A., et al.
Published: (2026)
by: Grossman, Thomas A., et al.
Published: (2026)
RepoTransBench: A Real-World Multilingual Benchmark for Repository-Level Code Translation
by: Wang, Yanli, et al.
Published: (2024)
by: Wang, Yanli, et al.
Published: (2024)
From Completion to Editing: Unlocking Context-Aware Code Infilling via Search-and-Replace Instruction Tuning
by: Zhang, Jiajun, et al.
Published: (2026)
by: Zhang, Jiajun, et al.
Published: (2026)
Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks
by: Zhao, Songwen, et al.
Published: (2025)
by: Zhao, Songwen, et al.
Published: (2025)
EffiBench: Benchmarking the Efficiency of Automatically Generated Code
by: Huang, Dong, et al.
Published: (2024)
by: Huang, Dong, et al.
Published: (2024)
Leveraging Test Driven Development with Large Language Models for Reliable and Verifiable Spreadsheet Code Generation: A Research Framework
by: Thorne, Simon, et al.
Published: (2025)
by: Thorne, Simon, et al.
Published: (2025)
TableTalk: Scaffolding Spreadsheet Development with a Language Agent
by: Liang, Jenny T., et al.
Published: (2025)
by: Liang, Jenny T., et al.
Published: (2025)
BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
by: Zhang, Zehua, et al.
Published: (2025)
by: Zhang, Zehua, et al.
Published: (2025)
CodeFlowBench: A Multi-turn, Iterative Benchmark for Complex Code Generation
by: Wang, Sizhe, et al.
Published: (2025)
by: Wang, Sizhe, et al.
Published: (2025)
Decompile-Bench: Million-Scale Binary-Source Function Pairs for Real-World Binary Decompilation
by: Tan, Hanzhuo, et al.
Published: (2025)
by: Tan, Hanzhuo, et al.
Published: (2025)
DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
by: Liu, Zhou, et al.
Published: (2025)
by: Liu, Zhou, et al.
Published: (2025)
SoAy: A Solution-based LLM API-using Methodology for Academic Information Seeking
by: Wang, Yuanchun, et al.
Published: (2024)
by: Wang, Yuanchun, et al.
Published: (2024)
Recommended Practices for Spreadsheet Testing
by: Panko, Raymond R.
Published: (2007)
by: Panko, Raymond R.
Published: (2007)
TableLLM: Enabling Tabular Data Manipulation by LLMs in Real Office Usage Scenarios
by: Zhang, Xiaokang, et al.
Published: (2024)
by: Zhang, Xiaokang, et al.
Published: (2024)
FireBench: Evaluating Instruction Following in Enterprise and API-Driven LLM Applications
by: Zhang, Yunfan, et al.
Published: (2026)
by: Zhang, Yunfan, et al.
Published: (2026)
CR-Bench: Evaluating the Real-World Utility of AI Code Review Agents
by: Pereira, Kristen, et al.
Published: (2026)
by: Pereira, Kristen, et al.
Published: (2026)
FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization
by: Jing, Huihao, et al.
Published: (2026)
by: Jing, Huihao, et al.
Published: (2026)
CoreCodeBench: Decoupling Code Intelligence via Fine-Grained Repository-Level Tasks
by: Fu, Lingyue, et al.
Published: (2025)
by: Fu, Lingyue, et al.
Published: (2025)
AutoBench: Automatic Testbench Generation and Evaluation Using LLMs for HDL Design
by: Qiu, Ruidi, et al.
Published: (2024)
by: Qiu, Ruidi, et al.
Published: (2024)
ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation
by: Zhang, Chenchen, et al.
Published: (2025)
by: Zhang, Chenchen, et al.
Published: (2025)
Model Editing for LLMs4Code: How Far are We?
by: Li, Xiaopeng, et al.
Published: (2024)
by: Li, Xiaopeng, et al.
Published: (2024)
RealBench: A Repo-Level Code Generation Benchmark Aligned with Real-World Software Development Practices
by: Li, Jia, et al.
Published: (2026)
by: Li, Jia, et al.
Published: (2026)
Firefly: Illuminating Large-Scale Verified Tool-Call Data Generation from Real APIs
by: Lu, Yuxuan, et al.
Published: (2026)
by: Lu, Yuxuan, et al.
Published: (2026)
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
by: Chou, Jason, et al.
Published: (2025)
by: Chou, Jason, et al.
Published: (2025)
Similar Items
-
Dynamic Scaling of Unit Tests for Code Reward Modeling
by: Ma, Zeyao, et al.
Published: (2025) -
Classification of Spreadsheet Errors
by: Rajalingham, Kamalasen, et al.
Published: (2008) -
Drivers of the Cost of Spreadsheet Audit
by: Colver, David
Published: (2011) -
Consensus-Free Spreadsheet Integration
by: Baylor, Brandon, et al.
Published: (2022) -
SODBench: A Large Language Model Approach to Documenting Spreadsheet Operations
by: Indika, Amila, et al.
Published: (2025)