JBE-QA: Japanese Bar Exam QA Dataset for Assessing Legal Domain Knowledge

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cao, Zhihan, Nishino, Fumihito, Yamada, Hiroaki, Thanh, Nguyen Ha, Miyao, Yusuke, Satoh, Ken
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915643109933056
author Cao, Zhihan
Nishino, Fumihito
Yamada, Hiroaki
Thanh, Nguyen Ha
Miyao, Yusuke
Satoh, Ken
author_facet Cao, Zhihan
Nishino, Fumihito
Yamada, Hiroaki
Thanh, Nguyen Ha
Miyao, Yusuke
Satoh, Ken
contents We introduce JBE-QA, a Japanese Bar Exam Question-Answering dataset to evaluate large language models' legal knowledge. Derived from the multiple-choice (tanto-shiki) section of the Japanese bar exam (2015-2024), JBE-QA provides the first comprehensive benchmark for Japanese legal-domain evaluation of LLMs. It covers the Civil Code, the Penal Code, and the Constitution, extending beyond the Civil Code focus of prior Japanese resources. Each question is decomposed into independent true/false judgments with structured contextual fields. The dataset contains 3,464 items with balanced labels. We evaluate 26 LLMs, including proprietary, open-weight, Japanese-specialised, and reasoning models. Our results show that proprietary models with reasoning enabled perform best, and the Constitution questions are generally easier than the Civil Code or the Penal Code questions.
format Preprint
id arxiv_https___arxiv_org_abs_2511_22869
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle JBE-QA: Japanese Bar Exam QA Dataset for Assessing Legal Domain Knowledge
Cao, Zhihan
Nishino, Fumihito
Yamada, Hiroaki
Thanh, Nguyen Ha
Miyao, Yusuke
Satoh, Ken
Computation and Language
We introduce JBE-QA, a Japanese Bar Exam Question-Answering dataset to evaluate large language models' legal knowledge. Derived from the multiple-choice (tanto-shiki) section of the Japanese bar exam (2015-2024), JBE-QA provides the first comprehensive benchmark for Japanese legal-domain evaluation of LLMs. It covers the Civil Code, the Penal Code, and the Constitution, extending beyond the Civil Code focus of prior Japanese resources. Each question is decomposed into independent true/false judgments with structured contextual fields. The dataset contains 3,464 items with balanced labels. We evaluate 26 LLMs, including proprietary, open-weight, Japanese-specialised, and reasoning models. Our results show that proprietary models with reasoning enabled perform best, and the Constitution questions are generally easier than the Civil Code or the Penal Code questions.
title JBE-QA: Japanese Bar Exam QA Dataset for Assessing Legal Domain Knowledge
topic Computation and Language
url https://arxiv.org/abs/2511.22869