OpenExempt: A Diagnostic Benchmark for Legal Reasoning and a Framework for Creating Custom Benchmarks on Demand
Fuente:
arXiv
Salvato in:
| Autori principali: | Servantez, Sergio, Lawsky, Sarah B., Jain, Rajiv, Linna Jr., Daniel W., Hammond, Kristian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Chain of Logic: Rule-Based Reasoning with Large Language Models
di: Servantez, Sergio, et al.
Pubblicazione: (2024)
di: Servantez, Sergio, et al.
Pubblicazione: (2024)
A Reasoning-Focused Legal Retrieval Benchmark
di: Zheng, Lucia, et al.
Pubblicazione: (2025)
di: Zheng, Lucia, et al.
Pubblicazione: (2025)
LegalRikai: Open Benchmark -- Benchmark for Complex Japanese Corporate Legal Tasks
di: Fujita, Shogo, et al.
Pubblicazione: (2025)
di: Fujita, Shogo, et al.
Pubblicazione: (2025)
ALARB: An Arabic Legal Argument Reasoning Benchmark
di: Shairah, Harethah Abu, et al.
Pubblicazione: (2025)
di: Shairah, Harethah Abu, et al.
Pubblicazione: (2025)
Korean Canonical Legal Benchmark: Toward Knowledge-Independent Evaluation of LLMs' Legal Reasoning Capabilities
di: Oh, Hongseok, et al.
Pubblicazione: (2025)
di: Oh, Hongseok, et al.
Pubblicazione: (2025)
Structured Definitions and Segmentations for Legal Reasoning in LLMs: A Study on Indian Legal Data
di: Khatri, Mann, et al.
Pubblicazione: (2025)
di: Khatri, Mann, et al.
Pubblicazione: (2025)
Challenges for Generative AI in Legal Reasoning
di: Linna, Eljas, et al.
Pubblicazione: (2025)
di: Linna, Eljas, et al.
Pubblicazione: (2025)
LEXam: Benchmarking Legal Reasoning on 340 Law Exams
di: Fan, Yu, et al.
Pubblicazione: (2025)
di: Fan, Yu, et al.
Pubblicazione: (2025)
MAWARITH: A Dataset and Benchmark for Legal Inheritance Reasoning with LLMs
di: Bouchekif, Abdessalam, et al.
Pubblicazione: (2026)
di: Bouchekif, Abdessalam, et al.
Pubblicazione: (2026)
AR-BENCH: Benchmarking Legal Reasoning with Judgment Error Detection, Classification and Correction
di: Li, Yifei, et al.
Pubblicazione: (2026)
di: Li, Yifei, et al.
Pubblicazione: (2026)
IL-TUR: Benchmark for Indian Legal Text Understanding and Reasoning
di: Joshi, Abhinav, et al.
Pubblicazione: (2024)
di: Joshi, Abhinav, et al.
Pubblicazione: (2024)
MASLegalBench: Benchmarking Multi-Agent Systems in Deductive Legal Reasoning
di: Jing, Huihao, et al.
Pubblicazione: (2025)
di: Jing, Huihao, et al.
Pubblicazione: (2025)
CSCBench: A PVC Diagnostic Benchmark for Commodity Supply Chain Reasoning
di: Cui, Yaxin, et al.
Pubblicazione: (2026)
di: Cui, Yaxin, et al.
Pubblicazione: (2026)
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models
di: Zhu, Yakun, et al.
Pubblicazione: (2025)
di: Zhu, Yakun, et al.
Pubblicazione: (2025)
OKBench: Democratizing LLM Evaluation with Fully Automated, On-Demand, Open Knowledge Benchmarking
di: Li, Yanhong, et al.
Pubblicazione: (2025)
di: Li, Yanhong, et al.
Pubblicazione: (2025)
Legal-DC: Benchmarking Retrieval-Augmented Generation for Legal Documents
di: Li, Yaocong, et al.
Pubblicazione: (2026)
di: Li, Yaocong, et al.
Pubblicazione: (2026)
Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases
di: AlDahoul, Nouar, et al.
Pubblicazione: (2025)
di: AlDahoul, Nouar, et al.
Pubblicazione: (2025)
Your Reasoning Benchmark May Not Test Reasoning: Revealing Perception Bottleneck in Abstract Reasoning Benchmarks
di: Wang, Xinhe, et al.
Pubblicazione: (2025)
di: Wang, Xinhe, et al.
Pubblicazione: (2025)
Benchmarking Legal RAG: The Promise and Limits of AI Statutory Surveys
di: Afane, Mohamed, et al.
Pubblicazione: (2026)
di: Afane, Mohamed, et al.
Pubblicazione: (2026)
GreekBarBench: A Challenging Benchmark for Free-Text Legal Reasoning and Citations
di: Chlapanis, Odysseas S., et al.
Pubblicazione: (2025)
di: Chlapanis, Odysseas S., et al.
Pubblicazione: (2025)
BriefMe: A Legal NLP Benchmark for Assisting with Legal Briefs
di: Woo, Jesse, et al.
Pubblicazione: (2025)
di: Woo, Jesse, et al.
Pubblicazione: (2025)
LegalCore: A Dataset for Event Coreference Resolution in Legal Documents
di: Wei, Kangda, et al.
Pubblicazione: (2025)
di: Wei, Kangda, et al.
Pubblicazione: (2025)
A Benchmark Dataset and Evaluation Framework for Vietnamese Large Language Models in Customer Support
di: Nguyen, Long S. T., et al.
Pubblicazione: (2025)
di: Nguyen, Long S. T., et al.
Pubblicazione: (2025)
One Law, Many Languages: Benchmarking Multilingual Legal Reasoning for Judicial Support
di: Stern, Ronja, et al.
Pubblicazione: (2023)
di: Stern, Ronja, et al.
Pubblicazione: (2023)
BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law
di: Nagl, Sebastian, et al.
Pubblicazione: (2026)
di: Nagl, Sebastian, et al.
Pubblicazione: (2026)
VLegal-Bench: Cognitively Grounded Benchmark for Vietnamese Legal Reasoning of Large Language Models
di: Dong, Nguyen Tien, et al.
Pubblicazione: (2025)
di: Dong, Nguyen Tien, et al.
Pubblicazione: (2025)
UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning
di: Ovcharov, Volodymyr
Pubblicazione: (2026)
di: Ovcharov, Volodymyr
Pubblicazione: (2026)
Benchmarking and Learning Real-World Customer Service Dialogue
di: Gao, Tianhong, et al.
Pubblicazione: (2025)
di: Gao, Tianhong, et al.
Pubblicazione: (2025)
PARSE: An Open-Domain Reasoning Question Answering Benchmark for Persian
di: Mozafari, Jamshid, et al.
Pubblicazione: (2026)
di: Mozafari, Jamshid, et al.
Pubblicazione: (2026)
Open Shouldn't Mean Exempt: Open-Source Exceptionalism and Generative AI
di: Atkinson, David
Pubblicazione: (2025)
di: Atkinson, David
Pubblicazione: (2025)
The Massive Legal Embedding Benchmark (MLEB)
di: Butler, Umar, et al.
Pubblicazione: (2025)
di: Butler, Umar, et al.
Pubblicazione: (2025)
LegalBench.PT: A Benchmark for Portuguese Law
di: Canaverde, Beatriz, et al.
Pubblicazione: (2025)
di: Canaverde, Beatriz, et al.
Pubblicazione: (2025)
PILOT-Bench: A Benchmark for Legal Reasoning in the Patent Domain with IRAC-Aligned Classification Tasks
di: Jang, Yehoon, et al.
Pubblicazione: (2026)
di: Jang, Yehoon, et al.
Pubblicazione: (2026)
RingSQL: Generating Synthetic Data with Schema-Independent Templates for Text-to-SQL Reasoning Models
di: Sterbentz, Marko, et al.
Pubblicazione: (2026)
di: Sterbentz, Marko, et al.
Pubblicazione: (2026)
CharacterBench: Benchmarking Character Customization of Large Language Models
di: Zhou, Jinfeng, et al.
Pubblicazione: (2024)
di: Zhou, Jinfeng, et al.
Pubblicazione: (2024)
CLaw: Benchmarking Chinese Legal Knowledge in Large Language Models - A Fine-grained Corpus and Reasoning Analysis
di: Xu, Xinzhe, et al.
Pubblicazione: (2025)
di: Xu, Xinzhe, et al.
Pubblicazione: (2025)
OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMs
di: Wang, Yuxia, et al.
Pubblicazione: (2024)
di: Wang, Yuxia, et al.
Pubblicazione: (2024)
Beyond Chains of Thought: Benchmarking Latent-Space Reasoning Abilities in Large Language Models
di: Hagendorff, Thilo, et al.
Pubblicazione: (2025)
di: Hagendorff, Thilo, et al.
Pubblicazione: (2025)
LexTime: A Benchmark for Temporal Ordering of Legal Events
di: Barale, Claire, et al.
Pubblicazione: (2025)
di: Barale, Claire, et al.
Pubblicazione: (2025)
CodeBenchGen: Creating Scalable Execution-based Code Generation Benchmarks
di: Xie, Yiqing, et al.
Pubblicazione: (2024)
di: Xie, Yiqing, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Chain of Logic: Rule-Based Reasoning with Large Language Models
di: Servantez, Sergio, et al.
Pubblicazione: (2024) -
A Reasoning-Focused Legal Retrieval Benchmark
di: Zheng, Lucia, et al.
Pubblicazione: (2025) -
LegalRikai: Open Benchmark -- Benchmark for Complex Japanese Corporate Legal Tasks
di: Fujita, Shogo, et al.
Pubblicazione: (2025) -
ALARB: An Arabic Legal Argument Reasoning Benchmark
di: Shairah, Harethah Abu, et al.
Pubblicazione: (2025) -
Korean Canonical Legal Benchmark: Toward Knowledge-Independent Evaluation of LLMs' Legal Reasoning Capabilities
di: Oh, Hongseok, et al.
Pubblicazione: (2025)