Expert Evaluation of LLM's Open-Ended Legal Reasoning on the Japanese Bar Exam Writing Task
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Choi, Jungmin, Sakaguchi, Keisuke, Yamada, Hiroaki |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
JADE: Expert-Grounded Dynamic Evaluation for Open-Ended Professional Tasks
von: Lin, Lanbo, et al.
Veröffentlicht: (2026)
von: Lin, Lanbo, et al.
Veröffentlicht: (2026)
Japanese Tort-case Dataset for Rationale-supported Legal Judgment Prediction
von: Yamada, Hiroaki, et al.
Veröffentlicht: (2023)
von: Yamada, Hiroaki, et al.
Veröffentlicht: (2023)
R2-Write: Reflection and Revision for Open-Ended Writing with Deep Reasoning
von: Liu, Wanlong, et al.
Veröffentlicht: (2026)
von: Liu, Wanlong, et al.
Veröffentlicht: (2026)
A Llama walks into the 'Bar': Efficient Supervised Fine-Tuning for Legal Reasoning in the Multi-state Bar Exam
von: Fernandes, Rean, et al.
Veröffentlicht: (2025)
von: Fernandes, Rean, et al.
Veröffentlicht: (2025)
AHP-Powered LLM Reasoning for Multi-Criteria Evaluation of Open-Ended Responses
von: Lu, Xiaotian, et al.
Veröffentlicht: (2024)
von: Lu, Xiaotian, et al.
Veröffentlicht: (2024)
JBE-QA: Japanese Bar Exam QA Dataset for Assessing Legal Domain Knowledge
von: Cao, Zhihan, et al.
Veröffentlicht: (2025)
von: Cao, Zhihan, et al.
Veröffentlicht: (2025)
A Mixture-of-Experts Approach to Few-Shot Task Transfer in Open-Ended Text Worlds
von: Cui, Christopher Z., et al.
Veröffentlicht: (2024)
von: Cui, Christopher Z., et al.
Veröffentlicht: (2024)
Self-Training Meets Consistency: Improving LLMs' Reasoning with Consistency-Driven Rationale Evaluation
von: Lee, Jaehyeok, et al.
Veröffentlicht: (2024)
von: Lee, Jaehyeok, et al.
Veröffentlicht: (2024)
ModeX: Evaluator-Free Best-of-N Selection for Open-Ended Generation
von: Choi, Hyeong Kyu, et al.
Veröffentlicht: (2026)
von: Choi, Hyeong Kyu, et al.
Veröffentlicht: (2026)
LLM Agents Beyond Utility: An Open-Ended Perspective
von: Nachkov, Asen, et al.
Veröffentlicht: (2025)
von: Nachkov, Asen, et al.
Veröffentlicht: (2025)
Reverse-Engineered Reasoning for Open-Ended Generation
von: Wang, Haozhe, et al.
Veröffentlicht: (2025)
von: Wang, Haozhe, et al.
Veröffentlicht: (2025)
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation
von: Tran, Khanh-Tung, et al.
Veröffentlicht: (2025)
von: Tran, Khanh-Tung, et al.
Veröffentlicht: (2025)
GPTs and Language Barrier: A Cross-Lingual Legal QA Examination
von: Nguyen, Ha-Thanh, et al.
Veröffentlicht: (2024)
von: Nguyen, Ha-Thanh, et al.
Veröffentlicht: (2024)
On Creativity and Open-Endedness
von: Soros, L. B., et al.
Veröffentlicht: (2024)
von: Soros, L. B., et al.
Veröffentlicht: (2024)
Unlocking Prompt Infilling Capability for Diffusion Language Models
von: Fujinuma, Yoshinari, et al.
Veröffentlicht: (2026)
von: Fujinuma, Yoshinari, et al.
Veröffentlicht: (2026)
LeMAJ (Legal LLM-as-a-Judge): Bridging Legal Reasoning and LLM Evaluation
von: Enguehard, Joseph, et al.
Veröffentlicht: (2025)
von: Enguehard, Joseph, et al.
Veröffentlicht: (2025)
LegalScore: Development of a Benchmark for Evaluating AI Models in Legal Career Exams in Brazil
von: Caparroz, Roberto, et al.
Veröffentlicht: (2025)
von: Caparroz, Roberto, et al.
Veröffentlicht: (2025)
Extending RLVR to Open-Ended Tasks via Verifiable Multiple-Choice Reformulation
von: Zhang, Mengyu, et al.
Veröffentlicht: (2025)
von: Zhang, Mengyu, et al.
Veröffentlicht: (2025)
Automatic Legal Writing Evaluation of LLMs
von: Pires, Ramon, et al.
Veröffentlicht: (2025)
von: Pires, Ramon, et al.
Veröffentlicht: (2025)
LEXam: Benchmarking Legal Reasoning on 340 Law Exams
von: Fan, Yu, et al.
Veröffentlicht: (2025)
von: Fan, Yu, et al.
Veröffentlicht: (2025)
Performance Evaluation of Open-Source Large Language Models for Assisting Pathology Report Writing in Japanese
von: Kawai, Masataka, et al.
Veröffentlicht: (2026)
von: Kawai, Masataka, et al.
Veröffentlicht: (2026)
Do LLMs Exhibit Human-Like Reasoning? Evaluating Theory of Mind in LLMs for Open-Ended Responses
von: Amirizaniani, Maryam, et al.
Veröffentlicht: (2024)
von: Amirizaniani, Maryam, et al.
Veröffentlicht: (2024)
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
von: Xie, Tianbao, et al.
Veröffentlicht: (2024)
von: Xie, Tianbao, et al.
Veröffentlicht: (2024)
Generation Space Size: Understanding and Calibrating Open-Endedness of LLM Generations
von: Yu, Sunny, et al.
Veröffentlicht: (2025)
von: Yu, Sunny, et al.
Veröffentlicht: (2025)
Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering
von: Choi, Changin, et al.
Veröffentlicht: (2025)
von: Choi, Changin, et al.
Veröffentlicht: (2025)
Multi-News+: Cost-efficient Dataset Cleansing via LLM-based Data Annotation
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
Pessimistic Verification for Open Ended Math Questions
von: Huang, Yanxing, et al.
Veröffentlicht: (2025)
von: Huang, Yanxing, et al.
Veröffentlicht: (2025)
Embodied World Models Emerge from Navigational Task in Open-Ended Environments
von: Jin, Li, et al.
Veröffentlicht: (2025)
von: Jin, Li, et al.
Veröffentlicht: (2025)
Open-Ended Task Discovery via Bayesian Optimization
von: Adachi, Masaki, et al.
Veröffentlicht: (2026)
von: Adachi, Masaki, et al.
Veröffentlicht: (2026)
Tru-POMDP: Task Planning Under Uncertainty via Tree of Hypotheses and Open-Ended POMDPs
von: Tang, Wenjing, et al.
Veröffentlicht: (2025)
von: Tang, Wenjing, et al.
Veröffentlicht: (2025)
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
von: Lee, Jaehyeok, et al.
Veröffentlicht: (2026)
von: Lee, Jaehyeok, et al.
Veröffentlicht: (2026)
AI Gamestore: Scalable, Open-Ended Evaluation of Machine General Intelligence with Human Games
von: Ying, Lance, et al.
Veröffentlicht: (2026)
von: Ying, Lance, et al.
Veröffentlicht: (2026)
Assessing the Reliability of Large Language Models in the Bengali Legal Context: A Comparative Evaluation Using LLM-as-Judge and Legal Experts
von: Aftahee, Sabik, et al.
Veröffentlicht: (2025)
von: Aftahee, Sabik, et al.
Veröffentlicht: (2025)
Assessing Bias in Metric Models for LLM Open-Ended Generation Bias Benchmarks
von: Demchak, Nathaniel, et al.
Veröffentlicht: (2024)
von: Demchak, Nathaniel, et al.
Veröffentlicht: (2024)
PILOT-Bench: A Benchmark for Legal Reasoning in the Patent Domain with IRAC-Aligned Classification Tasks
von: Jang, Yehoon, et al.
Veröffentlicht: (2026)
von: Jang, Yehoon, et al.
Veröffentlicht: (2026)
Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning
von: Ye, Zhiling, et al.
Veröffentlicht: (2025)
von: Ye, Zhiling, et al.
Veröffentlicht: (2025)
Open-Ended Multi-Modal Relational Reasoning for Video Question Answering
von: Luo, Haozheng, et al.
Veröffentlicht: (2020)
von: Luo, Haozheng, et al.
Veröffentlicht: (2020)
LLM-jp: A Cross-organizational Project for the Research and Development of Fully Open Japanese LLMs
von: LLM-jp, et al.
Veröffentlicht: (2024)
von: LLM-jp, et al.
Veröffentlicht: (2024)
DS-STAR: Data Science Agent for Solving Diverse Tasks across Heterogeneous Formats and Open-Ended Queries
von: Nam, Jaehyun, et al.
Veröffentlicht: (2025)
von: Nam, Jaehyun, et al.
Veröffentlicht: (2025)
Safety Must Precede the Deployment of Open-Ended AI
von: Sheth, Ivaxi, et al.
Veröffentlicht: (2025)
von: Sheth, Ivaxi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
JADE: Expert-Grounded Dynamic Evaluation for Open-Ended Professional Tasks
von: Lin, Lanbo, et al.
Veröffentlicht: (2026) -
Japanese Tort-case Dataset for Rationale-supported Legal Judgment Prediction
von: Yamada, Hiroaki, et al.
Veröffentlicht: (2023) -
R2-Write: Reflection and Revision for Open-Ended Writing with Deep Reasoning
von: Liu, Wanlong, et al.
Veröffentlicht: (2026) -
A Llama walks into the 'Bar': Efficient Supervised Fine-Tuning for Legal Reasoning in the Multi-state Bar Exam
von: Fernandes, Rean, et al.
Veröffentlicht: (2025) -
AHP-Powered LLM Reasoning for Multi-Criteria Evaluation of Open-Ended Responses
von: Lu, Xiaotian, et al.
Veröffentlicht: (2024)