LegalRikai: Open Benchmark -- Benchmark for Complex Japanese Corporate Legal Tasks
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Fujita, Shogo, Naraki, Yuji, Zhu, Yiqing, Mori, Shinsuke |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Legal-DC: Benchmarking Retrieval-Augmented Generation for Legal Documents
par: Li, Yaocong, et autres
Publié: (2026)
par: Li, Yaocong, et autres
Publié: (2026)
BriefMe: A Legal NLP Benchmark for Assisting with Legal Briefs
par: Woo, Jesse, et autres
Publié: (2025)
par: Woo, Jesse, et autres
Publié: (2025)
OpenExempt: A Diagnostic Benchmark for Legal Reasoning and a Framework for Creating Custom Benchmarks on Demand
par: Servantez, Sergio, et autres
Publié: (2026)
par: Servantez, Sergio, et autres
Publié: (2026)
LexSumm and LexT5: Benchmarking and Modeling Legal Summarization Tasks in English
par: Santosh, T. Y. S. S., et autres
Publié: (2024)
par: Santosh, T. Y. S. S., et autres
Publié: (2024)
A Reasoning-Focused Legal Retrieval Benchmark
par: Zheng, Lucia, et autres
Publié: (2025)
par: Zheng, Lucia, et autres
Publié: (2025)
Korean Canonical Legal Benchmark: Toward Knowledge-Independent Evaluation of LLMs' Legal Reasoning Capabilities
par: Oh, Hongseok, et autres
Publié: (2025)
par: Oh, Hongseok, et autres
Publié: (2025)
LegalBench.PT: A Benchmark for Portuguese Law
par: Canaverde, Beatriz, et autres
Publié: (2025)
par: Canaverde, Beatriz, et autres
Publié: (2025)
The Massive Legal Embedding Benchmark (MLEB)
par: Butler, Umar, et autres
Publié: (2025)
par: Butler, Umar, et autres
Publié: (2025)
LegalEval-Q: A New Benchmark for The Quality Evaluation of LLM-Generated Legal Text
par: yunhan, Li, et autres
Publié: (2025)
par: yunhan, Li, et autres
Publié: (2025)
LexTime: A Benchmark for Temporal Ordering of Legal Events
par: Barale, Claire, et autres
Publié: (2025)
par: Barale, Claire, et autres
Publié: (2025)
Benchmarking Legal RAG: The Promise and Limits of AI Statutory Surveys
par: Afane, Mohamed, et autres
Publié: (2026)
par: Afane, Mohamed, et autres
Publié: (2026)
ArabLegalEval: A Multitask Benchmark for Assessing Arabic Legal Knowledge in Large Language Models
par: Hijazi, Faris, et autres
Publié: (2024)
par: Hijazi, Faris, et autres
Publié: (2024)
ALARB: An Arabic Legal Argument Reasoning Benchmark
par: Shairah, Harethah Abu, et autres
Publié: (2025)
par: Shairah, Harethah Abu, et autres
Publié: (2025)
Indian Legal NLP Benchmarks : A Survey
par: Kalamkar, Prathamesh, et autres
Publié: (2021)
par: Kalamkar, Prathamesh, et autres
Publié: (2021)
PILOT-Bench: A Benchmark for Legal Reasoning in the Patent Domain with IRAC-Aligned Classification Tasks
par: Jang, Yehoon, et autres
Publié: (2026)
par: Jang, Yehoon, et autres
Publié: (2026)
LEXTREME: A Multi-Lingual and Multi-Task Benchmark for the Legal Domain
par: Niklaus, Joel, et autres
Publié: (2023)
par: Niklaus, Joel, et autres
Publié: (2023)
LexRel: Benchmarking Legal Relation Extraction for Chinese Civil Cases
par: Cai, Yida, et autres
Publié: (2025)
par: Cai, Yida, et autres
Publié: (2025)
LAiW: A Chinese Legal Large Language Models Benchmark
par: Dai, Yongfu, et autres
Publié: (2023)
par: Dai, Yongfu, et autres
Publié: (2023)
AppellateGen: A Benchmark for Appellate Legal Judgment Generation
par: Yang, Hongkun, et autres
Publié: (2026)
par: Yang, Hongkun, et autres
Publié: (2026)
Vaporetto: Efficient Japanese Tokenization Based on Improved Pointwise Linear Classification
par: Akabe, Koichi, et autres
Publié: (2024)
par: Akabe, Koichi, et autres
Publié: (2024)
HalluDetect: Detecting, Mitigating, and Benchmarking Hallucinations in Conversational Systems in the Legal Domain
par: Anaokar, Spandan, et autres
Publié: (2025)
par: Anaokar, Spandan, et autres
Publié: (2025)
CaseGen: A Benchmark for Multi-Stage Legal Case Documents Generation
par: Li, Haitao, et autres
Publié: (2025)
par: Li, Haitao, et autres
Publié: (2025)
AR-BENCH: Benchmarking Legal Reasoning with Judgment Error Detection, Classification and Correction
par: Li, Yifei, et autres
Publié: (2026)
par: Li, Yifei, et autres
Publié: (2026)
MAWARITH: A Dataset and Benchmark for Legal Inheritance Reasoning with LLMs
par: Bouchekif, Abdessalam, et autres
Publié: (2026)
par: Bouchekif, Abdessalam, et autres
Publié: (2026)
Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases
par: AlDahoul, Nouar, et autres
Publié: (2025)
par: AlDahoul, Nouar, et autres
Publié: (2025)
LegalLens Shared Task 2024: Legal Violation Identification in Unstructured Text
par: Hagag, Ben, et autres
Publié: (2024)
par: Hagag, Ben, et autres
Publié: (2024)
MASLegalBench: Benchmarking Multi-Agent Systems in Deductive Legal Reasoning
par: Jing, Huihao, et autres
Publié: (2025)
par: Jing, Huihao, et autres
Publié: (2025)
JUÁ -- A Benchmark for Information Retrieval in Brazilian Legal Text Collections
par: Pereira, Jayr, et autres
Publié: (2026)
par: Pereira, Jayr, et autres
Publié: (2026)
CaseFacts: A Benchmark for Legal Fact-Checking and Precedent Retrieval
par: Putta, Akshith Reddy, et autres
Publié: (2026)
par: Putta, Akshith Reddy, et autres
Publié: (2026)
IL-TUR: Benchmark for Indian Legal Text Understanding and Reasoning
par: Joshi, Abhinav, et autres
Publié: (2024)
par: Joshi, Abhinav, et autres
Publié: (2024)
GreekBarBench: A Challenging Benchmark for Free-Text Legal Reasoning and Citations
par: Chlapanis, Odysseas S., et autres
Publié: (2025)
par: Chlapanis, Odysseas S., et autres
Publié: (2025)
LexGenius: An Expert-Level Benchmark for Large Language Models in Legal General Intelligence
par: Liu, Wenjin, et autres
Publié: (2025)
par: Liu, Wenjin, et autres
Publié: (2025)
LegalBench-BR: A Benchmark for Evaluating Large Language Models on Brazilian Legal Decision Classification
par: Neto, Pedro Barbosa de Carvalho
Publié: (2026)
par: Neto, Pedro Barbosa de Carvalho
Publié: (2026)
LexEval: A Comprehensive Chinese Legal Benchmark for Evaluating Large Language Models
par: Li, Haitao, et autres
Publié: (2024)
par: Li, Haitao, et autres
Publié: (2024)
SinhaLegal: A Benchmark Corpus for Information Extraction and Analysis in Sinhala Legislative Texts
par: Lasandi, Minduli, et autres
Publié: (2026)
par: Lasandi, Minduli, et autres
Publié: (2026)
UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning
par: Ovcharov, Volodymyr
Publié: (2026)
par: Ovcharov, Volodymyr
Publié: (2026)
SwiLTra-Bench: The Swiss Legal Translation Benchmark
par: Niklaus, Joel, et autres
Publié: (2025)
par: Niklaus, Joel, et autres
Publié: (2025)
LEXam: Benchmarking Legal Reasoning on 340 Law Exams
par: Fan, Yu, et autres
Publié: (2025)
par: Fan, Yu, et autres
Publié: (2025)
JuDGE: Benchmarking Judgment Document Generation for Chinese Legal System
par: Su, Weihang, et autres
Publié: (2025)
par: Su, Weihang, et autres
Publié: (2025)
Bridging National and International Legal Data: Two Projects Based on the Japanese Legal Standard XML Schema for Comparative Law Studies
par: Nakamura, Makoto
Publié: (2026)
par: Nakamura, Makoto
Publié: (2026)
Documents similaires
-
Legal-DC: Benchmarking Retrieval-Augmented Generation for Legal Documents
par: Li, Yaocong, et autres
Publié: (2026) -
BriefMe: A Legal NLP Benchmark for Assisting with Legal Briefs
par: Woo, Jesse, et autres
Publié: (2025) -
OpenExempt: A Diagnostic Benchmark for Legal Reasoning and a Framework for Creating Custom Benchmarks on Demand
par: Servantez, Sergio, et autres
Publié: (2026) -
LexSumm and LexT5: Benchmarking and Modeling Legal Summarization Tasks in English
par: Santosh, T. Y. S. S., et autres
Publié: (2024) -
A Reasoning-Focused Legal Retrieval Benchmark
par: Zheng, Lucia, et autres
Publié: (2025)