Evaluating Semantic and Syntactic Understanding in Large Language Models for Payroll Systems

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Maclean, Hendrika, Cakmak, Mert Can, Mohammed, Muzakkiruddin Ahmed, Mandalawi, Shames Al, Talburt, John
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912885229223936
author Maclean, Hendrika
Cakmak, Mert Can
Mohammed, Muzakkiruddin Ahmed
Mandalawi, Shames Al
Talburt, John
author_facet Maclean, Hendrika
Cakmak, Mert Can
Mohammed, Muzakkiruddin Ahmed
Mandalawi, Shames Al
Talburt, John
contents Large language models are now used daily for writing, search, and analysis, and their natural language understanding continues to improve. However, they remain unreliable on exact numerical calculation and on producing outputs that are straightforward to audit. We study synthetic payroll system as a focused, high-stakes example and evaluate whether models can understand a payroll schema, apply rules in the right order, and deliver cent-accurate results. Our experiments span a tiered dataset from basic to complex cases, a spectrum of prompts from minimal baselines to schema-guided and reasoning variants, and multiple model families including GPT, Claude, Perplexity, Grok and Gemini. Results indicate clear regimes where careful prompting is sufficient and regimes where explicit computation is required. The work offers a compact, reproducible framework and practical guidance for deploying LLMs in settings that demand both accuracy and assurance.
format Preprint
id arxiv_https___arxiv_org_abs_2601_18012
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Evaluating Semantic and Syntactic Understanding in Large Language Models for Payroll Systems
Maclean, Hendrika
Cakmak, Mert Can
Mohammed, Muzakkiruddin Ahmed
Mandalawi, Shames Al
Talburt, John
Computation and Language
Artificial Intelligence
Large language models are now used daily for writing, search, and analysis, and their natural language understanding continues to improve. However, they remain unreliable on exact numerical calculation and on producing outputs that are straightforward to audit. We study synthetic payroll system as a focused, high-stakes example and evaluate whether models can understand a payroll schema, apply rules in the right order, and deliver cent-accurate results. Our experiments span a tiered dataset from basic to complex cases, a spectrum of prompts from minimal baselines to schema-guided and reasoning variants, and multiple model families including GPT, Claude, Perplexity, Grok and Gemini. Results indicate clear regimes where careful prompting is sufficient and regimes where explicit computation is required. The work offers a compact, reproducible framework and practical guidance for deploying LLMs in settings that demand both accuracy and assurance.
title Evaluating Semantic and Syntactic Understanding in Large Language Models for Payroll Systems
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2601.18012