Taxation Perspectives from Large Language Models: A Case Study on Additional Tax Penalties

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Choi, Eunkyung, Suh, Youngjin, Lee, Siun, Oh, Hongseok, Kang, Juheon, Hur, Won, Park, Hun, Hwang, Wonseok
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915782356631552
author Choi, Eunkyung
Suh, Youngjin
Lee, Siun
Oh, Hongseok
Kang, Juheon
Hur, Won
Park, Hun
Hwang, Wonseok
author_facet Choi, Eunkyung
Suh, Youngjin
Lee, Siun
Oh, Hongseok
Kang, Juheon
Hur, Won
Park, Hun
Hwang, Wonseok
contents How capable are large language models (LLMs) in the domain of taxation? Although numerous studies have explored the legal domain, research dedicated to taxation remains scarce. Moreover, the datasets used in these studies are either simplified, failing to reflect the real-world complexities, or not released as open-source. To address this gap, we introduce PLAT, a new benchmark designed to assess the ability of LLMs to predict the legitimacy of additional tax penalties. PLAT comprises 300 examples: (1) 100 binary-choice questions, (2) 100 multiple-choice questions, and (3) 100 essay-type questions, all derived from 100 Korean court precedents. PLAT is constructed to evaluate not only LLMs' understanding of tax law but also their performance in legal cases that require complex reasoning beyond straightforward application of statutes. Our systematic experiments with multiple LLMs reveal that (1) their baseline capabilities are limited, especially in cases involving conflicting issues that require a comprehensive understanding (not only of the statutes but also of the taxpayer's circumstances), and (2) LLMs struggle particularly with the "AC" stages of "IRAC" even for advanced reasoning models like o3, which actively employ inference-time scaling. The dataset is publicly available at: https://huggingface.co/collections/sma1-rmarud/plat-predicting-the-legitimacy-of-punitive-additional-tax
format Preprint
id arxiv_https___arxiv_org_abs_2503_03444
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Taxation Perspectives from Large Language Models: A Case Study on Additional Tax Penalties
Choi, Eunkyung
Suh, Youngjin
Lee, Siun
Oh, Hongseok
Kang, Juheon
Hur, Won
Park, Hun
Hwang, Wonseok
Computation and Language
Artificial Intelligence
How capable are large language models (LLMs) in the domain of taxation? Although numerous studies have explored the legal domain, research dedicated to taxation remains scarce. Moreover, the datasets used in these studies are either simplified, failing to reflect the real-world complexities, or not released as open-source. To address this gap, we introduce PLAT, a new benchmark designed to assess the ability of LLMs to predict the legitimacy of additional tax penalties. PLAT comprises 300 examples: (1) 100 binary-choice questions, (2) 100 multiple-choice questions, and (3) 100 essay-type questions, all derived from 100 Korean court precedents. PLAT is constructed to evaluate not only LLMs' understanding of tax law but also their performance in legal cases that require complex reasoning beyond straightforward application of statutes. Our systematic experiments with multiple LLMs reveal that (1) their baseline capabilities are limited, especially in cases involving conflicting issues that require a comprehensive understanding (not only of the statutes but also of the taxpayer's circumstances), and (2) LLMs struggle particularly with the "AC" stages of "IRAC" even for advanced reasoning models like o3, which actively employ inference-time scaling. The dataset is publicly available at: https://huggingface.co/collections/sma1-rmarud/plat-predicting-the-legitimacy-of-punitive-additional-tax
title Taxation Perspectives from Large Language Models: A Case Study on Additional Tax Penalties
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2503.03444