JBE-QA: Japanese Bar Exam QA Dataset for Assessing Legal Domain Knowledge
Fuente:
arXiv
Saved in:
| Main Authors: | Cao, Zhihan, Nishino, Fumihito, Yamada, Hiroaki, Thanh, Nguyen Ha, Miyao, Yusuke, Satoh, Ken |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GPTs and Language Barrier: A Cross-Lingual Legal QA Examination
by: Nguyen, Ha-Thanh, et al.
Published: (2024)
by: Nguyen, Ha-Thanh, et al.
Published: (2024)
KRAG Framework for Enhancing LLMs in the Legal Domain
by: Thanh, Nguyen Ha, et al.
Published: (2024)
by: Thanh, Nguyen Ha, et al.
Published: (2024)
PYTHEN: A Flexible Framework for Legal Reasoning in Python
by: Nguyen, Ha-Thanh, et al.
Published: (2026)
by: Nguyen, Ha-Thanh, et al.
Published: (2026)
Data Augmented Pipeline for Legal Information Extraction and Reasoning
by: Phuong, Nguyen Minh, et al.
Published: (2026)
by: Phuong, Nguyen Minh, et al.
Published: (2026)
Balancing Exploration and Exploitation in LLM using Soft RLLF for Enhanced Negation Understanding
by: Nguyen, Ha-Thanh, et al.
Published: (2024)
by: Nguyen, Ha-Thanh, et al.
Published: (2024)
BioGraphletQA: Knowledge-Anchored Generation of Complex QA Datasets
by: Jonker, Richard A. A., et al.
Published: (2026)
by: Jonker, Richard A. A., et al.
Published: (2026)
Legal2LogicICL: Improving Generalization in Transforming Legal Cases to Logical Formulas via Diverse Few-Shot Learning
by: Xue, Jieying, et al.
Published: (2026)
by: Xue, Jieying, et al.
Published: (2026)
RJUA-QA: A Comprehensive QA Dataset for Urology
by: Lyu, Shiwei, et al.
Published: (2023)
by: Lyu, Shiwei, et al.
Published: (2023)
Misalignment of Semantic Relation Knowledge between WordNet and Human Intuition
by: Cao, Zhihan, et al.
Published: (2024)
by: Cao, Zhihan, et al.
Published: (2024)
A Comprehensive Evaluation of Semantic Relation Knowledge of Pretrained Language Models and Humans
by: Cao, Zhihan, et al.
Published: (2024)
by: Cao, Zhihan, et al.
Published: (2024)
SealQA: Raising the Bar for Reasoning in Search-Augmented Language Models
by: Pham, Thinh, et al.
Published: (2025)
by: Pham, Thinh, et al.
Published: (2025)
On the Distinctive Co-occurrence Characteristics of Antonymy
by: Cao, Zhihan, et al.
Published: (2025)
by: Cao, Zhihan, et al.
Published: (2025)
JDocQA: Japanese Document Question Answering Dataset for Generative Language Models
by: Onami, Eri, et al.
Published: (2024)
by: Onami, Eri, et al.
Published: (2024)
Expert Evaluation of LLM's Open-Ended Legal Reasoning on the Japanese Bar Exam Writing Task
by: Choi, Jungmin, et al.
Published: (2026)
by: Choi, Jungmin, et al.
Published: (2026)
Enhancing Legal Document Retrieval: A Multi-Phase Approach with Large Language Models
by: Nguyen, Hai-Long, et al.
Published: (2024)
by: Nguyen, Hai-Long, et al.
Published: (2024)
Japanese Tort-case Dataset for Rationale-supported Legal Judgment Prediction
by: Yamada, Hiroaki, et al.
Published: (2023)
by: Yamada, Hiroaki, et al.
Published: (2023)
DomainCQA: Crafting Knowledge-Intensive QA from Domain-Specific Charts
by: Lu, Yujing, et al.
Published: (2025)
by: Lu, Yujing, et al.
Published: (2025)
AirQA: A Comprehensive QA Dataset for AI Research with Instance-Level Evaluation
by: Huang, Tiancheng, et al.
Published: (2025)
by: Huang, Tiancheng, et al.
Published: (2025)
ExpertGenQA: Open-ended QA generation in Specialized Domains
by: Shahgir, Haz Sameen, et al.
Published: (2025)
by: Shahgir, Haz Sameen, et al.
Published: (2025)
Formal Reasoning for Intelligent QA Systems: A Case Study in the Educational Domain
by: Bui, Tuan, et al.
Published: (2025)
by: Bui, Tuan, et al.
Published: (2025)
KET-QA: A Dataset for Knowledge Enhanced Table Question Answering
by: Hu, Mengkang, et al.
Published: (2024)
by: Hu, Mengkang, et al.
Published: (2024)
FictionalQA: A Dataset for Studying Memorization and Knowledge Acquisition
by: Kirchenbauer, John, et al.
Published: (2025)
by: Kirchenbauer, John, et al.
Published: (2025)
HumMusQA: A Human-written Music Understanding QA Benchmark Dataset
by: Weck, Benno, et al.
Published: (2026)
by: Weck, Benno, et al.
Published: (2026)
Syn-QA2: Evaluating False Assumptions in Long-tail Questions with Synthetic QA Datasets
by: Daswani, Ashwin, et al.
Published: (2024)
by: Daswani, Ashwin, et al.
Published: (2024)
PeruMedQA: Benchmarking Large Language Models (LLMs) on Peruvian Medical Exams -- Dataset Construction and Evaluation
by: Carrillo-Larco, Rodrigo M., et al.
Published: (2025)
by: Carrillo-Larco, Rodrigo M., et al.
Published: (2025)
Layer-of-Thoughts Prompting (LoT): Leveraging LLM-Based Retrieval with Constraint Hierarchies
by: Fungwacharakorn, Wachara, et al.
Published: (2024)
by: Fungwacharakorn, Wachara, et al.
Published: (2024)
LinkQA: Synthesizing Diverse QA from Multiple Seeds Strongly Linked by Knowledge Points
by: Zhang, Xuemiao, et al.
Published: (2025)
by: Zhang, Xuemiao, et al.
Published: (2025)
CUS-QA: Local-Knowledge-Oriented Open-Ended Question Answering Dataset
by: Libovický, Jindřich, et al.
Published: (2025)
by: Libovický, Jindřich, et al.
Published: (2025)
ViQA-COVID: COVID-19 Machine Reading Comprehension Dataset for Vietnamese
by: Nguyen-Phung, Hai-Chung, et al.
Published: (2025)
by: Nguyen-Phung, Hai-Chung, et al.
Published: (2025)
ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly
by: Hasegawa, Kimihiro, et al.
Published: (2025)
by: Hasegawa, Kimihiro, et al.
Published: (2025)
PolQA: Polish Question Answering Dataset
by: Rybak, Piotr, et al.
Published: (2022)
by: Rybak, Piotr, et al.
Published: (2022)
StorySparkQA: Expert-Annotated QA Pairs with Real-World Knowledge for Children's Story-Based Learning
by: Chen, Jiaju, et al.
Published: (2023)
by: Chen, Jiaju, et al.
Published: (2023)
Exploring the Impact of Occupational Personas on Domain-Specific QA
by: Kang, Eojin, et al.
Published: (2025)
by: Kang, Eojin, et al.
Published: (2025)
SEC-QA: A Systematic Evaluation Corpus for Financial QA
by: Lai, Viet Dac, et al.
Published: (2024)
by: Lai, Viet Dac, et al.
Published: (2024)
QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA
by: Dineen, Jacob, et al.
Published: (2025)
by: Dineen, Jacob, et al.
Published: (2025)
No Dataset Needed for Downstream Knowledge Benchmarking: Response Dispersion Inversely Correlates with Accuracy on Domain-specific QA
by: Simione II, Robert L
Published: (2024)
by: Simione II, Robert L
Published: (2024)
RetrievalQA: Assessing Adaptive Retrieval-Augmented Generation for Short-form Open-Domain Question Answering
by: Zhang, Zihan, et al.
Published: (2024)
by: Zhang, Zihan, et al.
Published: (2024)
VlogQA: Task, Dataset, and Baseline Models for Vietnamese Spoken-Based Machine Reading Comprehension
by: Ngo, Thinh Phuoc, et al.
Published: (2024)
by: Ngo, Thinh Phuoc, et al.
Published: (2024)
CogRAG+: Cognitive-Level Guided Diagnosis and Remediation of Memory and Reasoning Deficiencies in Professional Exam QA
by: Wang, Xudong, et al.
Published: (2026)
by: Wang, Xudong, et al.
Published: (2026)
Contextual Breach: Assessing the Robustness of Transformer-based QA Models
by: Saadat, Asir, et al.
Published: (2024)
by: Saadat, Asir, et al.
Published: (2024)
Similar Items
-
GPTs and Language Barrier: A Cross-Lingual Legal QA Examination
by: Nguyen, Ha-Thanh, et al.
Published: (2024) -
KRAG Framework for Enhancing LLMs in the Legal Domain
by: Thanh, Nguyen Ha, et al.
Published: (2024) -
PYTHEN: A Flexible Framework for Legal Reasoning in Python
by: Nguyen, Ha-Thanh, et al.
Published: (2026) -
Data Augmented Pipeline for Legal Information Extraction and Reasoning
by: Phuong, Nguyen Minh, et al.
Published: (2026) -
Balancing Exploration and Exploitation in LLM using Soft RLLF for Enhanced Negation Understanding
by: Nguyen, Ha-Thanh, et al.
Published: (2024)