Assessing the Reliability of Large Language Models in the Bengali Legal Context: A Comparative Evaluation Using LLM-as-Judge and Legal Experts
Fuente:
arXiv
Saved in:
| Main Authors: | Aftahee, Sabik, Farhad, A. F. M., Mallik, Arpita, Dhar, Ratnajit, Karim, Jawadul, Noor, Nahiyan Bin, Solaiman, Ishmam Ahmed |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
823-OLT @ BUET DL Sprint 4.0: Context-Aware Windowing for ASR and Fine-Tuned Speaker Diarization in Bengali Long Form Audio
by: Dhar, Ratnajit, et al.
Published: (2026)
by: Dhar, Ratnajit, et al.
Published: (2026)
Integrating Machine Learning Ensembles and Large Language Models for Heart Disease Prediction Using Voting Fusion
by: Amin, Md. Tahsin, et al.
Published: (2026)
by: Amin, Md. Tahsin, et al.
Published: (2026)
LeMAJ (Legal LLM-as-a-Judge): Bridging Legal Reasoning and LLM Evaluation
by: Enguehard, Joseph, et al.
Published: (2025)
by: Enguehard, Joseph, et al.
Published: (2025)
LegalCiteBench: Evaluating Citation Reliability in Legal Language Models
by: Chen, Sijia, et al.
Published: (2026)
by: Chen, Sijia, et al.
Published: (2026)
A Comprehensive Framework for Reliable Legal AI: Combining Specialized Expert Systems and Adaptive Refinement
by: Nasir, Sidra, et al.
Published: (2024)
by: Nasir, Sidra, et al.
Published: (2024)
LegalAgentBench: Evaluating LLM Agents in Legal Domain
by: Li, Haitao, et al.
Published: (2024)
by: Li, Haitao, et al.
Published: (2024)
LLM-as-a-Judge: Rapid Evaluation of Legal Document Recommendation for Retrieval-Augmented Generation
by: Pradhan, Anu, et al.
Published: (2025)
by: Pradhan, Anu, et al.
Published: (2025)
Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools
by: Magesh, Varun, et al.
Published: (2024)
by: Magesh, Varun, et al.
Published: (2024)
The Judge Variable: Challenging Judge-Agnostic Legal Judgment Prediction
by: Zambrano, Guillaume
Published: (2025)
by: Zambrano, Guillaume
Published: (2025)
(A)I Am Not a Lawyer, But...: Engaging Legal Experts towards Responsible LLM Policies for Legal Advice
by: Cheong, Inyoung, et al.
Published: (2024)
by: Cheong, Inyoung, et al.
Published: (2024)
LegalOne: A Family of Foundation Models for Reliable Legal Reasoning
by: Li, Haitao, et al.
Published: (2026)
by: Li, Haitao, et al.
Published: (2026)
LegalWebAgent: Empowering Access to Justice via LLM-Based Web Agents
by: Tan, Jinzhe, et al.
Published: (2025)
by: Tan, Jinzhe, et al.
Published: (2025)
BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali
by: Adib, Shefayat E Shams, et al.
Published: (2026)
by: Adib, Shefayat E Shams, et al.
Published: (2026)
Expert Evaluation of LLM's Open-Ended Legal Reasoning on the Japanese Bar Exam Writing Task
by: Choi, Jungmin, et al.
Published: (2026)
by: Choi, Jungmin, et al.
Published: (2026)
LegalEval-Q: A New Benchmark for The Quality Evaluation of LLM-Generated Legal Text
by: yunhan, Li, et al.
Published: (2025)
by: yunhan, Li, et al.
Published: (2025)
Bone healing under different lay‐up configuration of carbon fiber‐reinforced PEEK composite plates
by: Agnieszka Sabik
Published: (2024)
by: Agnieszka Sabik
Published: (2024)
FedJudge: Federated Legal Large Language Model
by: Yue, Linan, et al.
Published: (2023)
by: Yue, Linan, et al.
Published: (2023)
Interface Design to Support Legal Reading and Writing: Insights from Interviews with Legal Experts
by: Swoopes, Chelse, et al.
Published: (2025)
by: Swoopes, Chelse, et al.
Published: (2025)
Retrieval Augmented Generation-based Large Language Models for Bridging Transportation Cybersecurity Legal Knowledge Gaps
by: Akbar, Khandakar Ashrafi, et al.
Published: (2025)
by: Akbar, Khandakar Ashrafi, et al.
Published: (2025)
Investigative Judges as a Legal Transplant: Finnish Nineteenth-Century Criminal Procedure in Comparative Perspective
by: Heikki Pihlajamäki
Published: (2021)
by: Heikki Pihlajamäki
Published: (2021)
LegalGraphRAG: Multi-Agent Graph Retrieval-Augmented Generation for Reliable Legal Reasoning
by: Chen, Zerui, et al.
Published: (2026)
by: Chen, Zerui, et al.
Published: (2026)
Exploiting LLM-as-a-Judge Disposition on Free Text Legal QA via Prompt Optimization
by: Elganayni, Mohamed Hesham, et al.
Published: (2026)
by: Elganayni, Mohamed Hesham, et al.
Published: (2026)
Evaluating Legal Reasoning Traces with Legal Issue Tree Rubrics
by: Lee, Jinu, et al.
Published: (2025)
by: Lee, Jinu, et al.
Published: (2025)
Judge Reliability Harness: Stress Testing the Reliability of LLM Judges
by: Dev, Sunishchal, et al.
Published: (2026)
by: Dev, Sunishchal, et al.
Published: (2026)
LegalCheck: Retrieval- and Context-Augmented Generation for Drafting Municipal Legal Advice Letters
by: van der Meer, Virgill, et al.
Published: (2026)
by: van der Meer, Virgill, et al.
Published: (2026)
Navigating Legal Frontiers in Robotic Healthcare: A Techno-Legal Appraisal in the Indian Context
by: Jadhav, Tanishka
Published: (2025)
by: Jadhav, Tanishka
Published: (2025)
VERT: Reliable LLM Judges for Radiology Report Evaluation
by: Bologna, Federica, et al.
Published: (2026)
by: Bologna, Federica, et al.
Published: (2026)
Evaluating the optimal duration of medication treatment for opioid use disorder
by: Corey J. Hayes, et al.
Published: (2026)
by: Corey J. Hayes, et al.
Published: (2026)
Foreign Investment in Cuba: Assessing the Legal Landscape
by: Melissa Johns
Published: (2003)
by: Melissa Johns
Published: (2003)
CALRK-Bench: Evaluating Context-Aware Legal Reasoning in Korean Law
by: Jung, JiHyeok, et al.
Published: (2026)
by: Jung, JiHyeok, et al.
Published: (2026)
ArabLegalEval: A Multitask Benchmark for Assessing Arabic Legal Knowledge in Large Language Models
by: Hijazi, Faris, et al.
Published: (2024)
by: Hijazi, Faris, et al.
Published: (2024)
When Safety Blocks Sense: Measuring Semantic Confusion in LLM Refusals
by: Anonto, Riad Ahmed, et al.
Published: (2025)
by: Anonto, Riad Ahmed, et al.
Published: (2025)
Place Matters: Comparing LLM Hallucination Rates for Place-Based Legal Queries
by: Curran, Damian, et al.
Published: (2025)
by: Curran, Damian, et al.
Published: (2025)
European Journal of Psychology Applied to Legal Context
Published: (2009)
Published: (2009)
The Legal Context of Education. Monograph Series 19.
by: Zuker, Marvin A.
Published: (1988)
by: Zuker, Marvin A.
Published: (1988)
Gavel: Agent Meets Checklist for Evaluating LLMs on Long-Context Legal Summarization
by: Dou, Yao, et al.
Published: (2026)
by: Dou, Yao, et al.
Published: (2026)
TathyaNyaya and FactLegalLlama: Advancing Factual Judgment Prediction and Explanation in the Indian Legal Context
by: Nigam, Shubham Kumar, et al.
Published: (2025)
by: Nigam, Shubham Kumar, et al.
Published: (2025)
JUDGEBERT: Assessing Legal Meaning Preservation Between Sentences
by: Beauchemin, David, et al.
Published: (2025)
by: Beauchemin, David, et al.
Published: (2025)
Capacity, Participation and Values in Comparative Legal Perspective
Published: (2023)
Published: (2023)
Bengali Document Layout Analysis -- A YOLOV8 Based Ensembling Approach
by: Ahmed, Nazmus Sakib, et al.
Published: (2023)
by: Ahmed, Nazmus Sakib, et al.
Published: (2023)
Similar Items
-
823-OLT @ BUET DL Sprint 4.0: Context-Aware Windowing for ASR and Fine-Tuned Speaker Diarization in Bengali Long Form Audio
by: Dhar, Ratnajit, et al.
Published: (2026) -
Integrating Machine Learning Ensembles and Large Language Models for Heart Disease Prediction Using Voting Fusion
by: Amin, Md. Tahsin, et al.
Published: (2026) -
LeMAJ (Legal LLM-as-a-Judge): Bridging Legal Reasoning and LLM Evaluation
by: Enguehard, Joseph, et al.
Published: (2025) -
LegalCiteBench: Evaluating Citation Reliability in Legal Language Models
by: Chen, Sijia, et al.
Published: (2026) -
A Comprehensive Framework for Reliable Legal AI: Combining Specialized Expert Systems and Adaptive Refinement
by: Nasir, Sidra, et al.
Published: (2024)