Measuring Faithfulness and Abstention: An Automated Pipeline for Evaluating LLM-Generated 3-ply Case-Based Legal Arguments
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Li, Gray, Morgan, Savelka, Jaromir, Ashley, Kevin D. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Retrieval-Based Multi-Label Legal Annotation: Extensible, Data-Efficient and Hallucination-Free
by: Zhang, Li, et al.
Published: (2026)
by: Zhang, Li, et al.
Published: (2026)
Do LLMs Truly Understand When a Precedent Is Overruled?
by: Zhang, Li, et al.
Published: (2025)
by: Zhang, Li, et al.
Published: (2025)
Mitigating Manipulation and Enhancing Persuasion: A Reflective Multi-Agent Approach for Legal Argument Generation
by: Zhang, Li, et al.
Published: (2025)
by: Zhang, Li, et al.
Published: (2025)
Thinking Longer, Not Always Smarter: Evaluating LLM Capabilities in Hierarchical Legal Reasoning
by: Zhang, Li, et al.
Published: (2025)
by: Zhang, Li, et al.
Published: (2025)
Using LLMs to Discover Legal Factors
by: Gray, Morgan, et al.
Published: (2024)
by: Gray, Morgan, et al.
Published: (2024)
Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation
by: Chen, Jiaju, et al.
Published: (2025)
by: Chen, Jiaju, et al.
Published: (2025)
RAG System for Supporting Japanese Litigation Procedures: Faithful Response Generation Complying with Legal Norms
by: Ishihara, Yuya, et al.
Published: (2025)
by: Ishihara, Yuya, et al.
Published: (2025)
One SPACE to Rule Them All: Jointly Mitigating Factuality and Faithfulness Hallucinations in LLMs
by: Wang, Pengbo, et al.
Published: (2025)
by: Wang, Pengbo, et al.
Published: (2025)
Japanese Tort-case Dataset for Rationale-supported Legal Judgment Prediction
by: Yamada, Hiroaki, et al.
Published: (2023)
by: Yamada, Hiroaki, et al.
Published: (2023)
OPENXRD: A Comprehensive Benchmark Framework for LLM/MLLM XRD Question Answering
by: Vosoughi, Ali, et al.
Published: (2025)
by: Vosoughi, Ali, et al.
Published: (2025)
LLM-Assisted Crisis Management: Building Advanced LLM Platforms for Effective Emergency Response and Public Collaboration
by: Otal, Hakan T., et al.
Published: (2024)
by: Otal, Hakan T., et al.
Published: (2024)
Crossing Linguistic Horizons: Finetuning and Comprehensive Evaluation of Vietnamese Large Language Models
by: Truong, Sang T., et al.
Published: (2024)
by: Truong, Sang T., et al.
Published: (2024)
Can Out-of-Distribution Evaluations Uncover Reliance on Shortcuts? A Case Study in Question Answering
by: Štefánik, Michal, et al.
Published: (2025)
by: Štefánik, Michal, et al.
Published: (2025)
Pretraining and Updates of Domain-Specific LLM: A Case Study in the Japanese Business Domain
by: Takahashi, Kosuke, et al.
Published: (2024)
by: Takahashi, Kosuke, et al.
Published: (2024)
LLM Evaluation Based on Aerospace Manufacturing Expertise: Automated Generation and Multi-Model Question Answering
by: Liu, Beiming, et al.
Published: (2025)
by: Liu, Beiming, et al.
Published: (2025)
LEGAL-UQA: A Low-Resource Urdu-English Dataset for Legal Question Answering
by: Faisal, Faizan, et al.
Published: (2024)
by: Faisal, Faizan, et al.
Published: (2024)
MultiLegalPile: A 689GB Multilingual Legal Corpus
by: Niklaus, Joel, et al.
Published: (2023)
by: Niklaus, Joel, et al.
Published: (2023)
Mediator: Memory-efficient LLM Merging with Less Parameter Conflicts and Uncertainty Based Routing
by: Lai, Kunfeng, et al.
Published: (2025)
by: Lai, Kunfeng, et al.
Published: (2025)
Curriculum Recommendations Using Transformer Base Model with InfoNCE Loss And Language Switching Method
by: Xu, Xiaonan, et al.
Published: (2024)
by: Xu, Xiaonan, et al.
Published: (2024)
Textual Data Bias Detection and Mitigation -- An Extensible Pipeline with Experimental Evaluation
by: Görge, Rebekka, et al.
Published: (2025)
by: Görge, Rebekka, et al.
Published: (2025)
Semantic Needles in Document Haystacks: Sensitivity Testing of LLM-as-a-Judge Similarity Scoring
by: Aksoy, Sinan G., et al.
Published: (2026)
by: Aksoy, Sinan G., et al.
Published: (2026)
Understanding the Effects of RLHF on the Quality and Detectability of LLM-Generated Texts
by: Xu, Beining, et al.
Published: (2025)
by: Xu, Beining, et al.
Published: (2025)
Forging GEMs: Advancing Greek NLP through Quality-Based Corpus Curation
by: Apostolopoulou, Alexandra, et al.
Published: (2025)
by: Apostolopoulou, Alexandra, et al.
Published: (2025)
LLMs as Deceptive Agents: How Role-Based Prompting Induces Semantic Ambiguity in Puzzle Tasks
by: Yoo, Seunghyun
Published: (2025)
by: Yoo, Seunghyun
Published: (2025)
FAIR-RAG: Faithful Adaptive Iterative Refinement for Retrieval-Augmented Generation
by: Asl, Mohammad Aghajani, et al.
Published: (2025)
by: Asl, Mohammad Aghajani, et al.
Published: (2025)
CATER: Leveraging LLM to Pioneer a Multidimensional, Reference-Independent Paradigm in Translation Quality Evaluation
by: IIDA, Kurando, et al.
Published: (2024)
by: IIDA, Kurando, et al.
Published: (2024)
HInter: Exposing Hidden Intersectional Bias in Large Language Models
by: Souani, Badr, et al.
Published: (2025)
by: Souani, Badr, et al.
Published: (2025)
Pay Attention to What You Need
by: Gao, Yifei, et al.
Published: (2023)
by: Gao, Yifei, et al.
Published: (2023)
LawInstruct: A Resource for Studying Language Model Adaptation to the Legal Domain
by: Niklaus, Joel, et al.
Published: (2024)
by: Niklaus, Joel, et al.
Published: (2024)
SwiLTra-Bench: The Swiss Legal Translation Benchmark
by: Niklaus, Joel, et al.
Published: (2025)
by: Niklaus, Joel, et al.
Published: (2025)
LEXam: Benchmarking Legal Reasoning on 340 Law Exams
by: Fan, Yu, et al.
Published: (2025)
by: Fan, Yu, et al.
Published: (2025)
FARSIQA: Faithful and Advanced RAG System for Islamic Question Answering
by: Asl, Mohammad Aghajani, et al.
Published: (2025)
by: Asl, Mohammad Aghajani, et al.
Published: (2025)
LEXTREME: A Multi-Lingual and Multi-Task Benchmark for the Legal Domain
by: Niklaus, Joel, et al.
Published: (2023)
by: Niklaus, Joel, et al.
Published: (2023)
Low-Resource Language Processing: An OCR-Driven Summarization and Translation Pipeline
by: Madhavi, Hrishit, et al.
Published: (2025)
by: Madhavi, Hrishit, et al.
Published: (2025)
One Law, Many Languages: Benchmarking Multilingual Legal Reasoning for Judicial Support
by: Stern, Ronja, et al.
Published: (2023)
by: Stern, Ronja, et al.
Published: (2023)
MetaCheckGPT -- A Multi-task Hallucination Detector Using LLM Uncertainty and Meta-models
by: Mehta, Rahul, et al.
Published: (2024)
by: Mehta, Rahul, et al.
Published: (2024)
Can AI Examine Novelty of Patents?: Novelty Evaluation Based on the Correspondence between Patent Claim and Prior Art
by: Ikoma, Hayato, et al.
Published: (2025)
by: Ikoma, Hayato, et al.
Published: (2025)
The Illusion of Role Separation: Hidden Shortcuts in LLM Role Learning (and How to Fix Them)
by: Wang, Zihao, et al.
Published: (2025)
by: Wang, Zihao, et al.
Published: (2025)
SVDq: 1.25-bit and 410x Key Cache Compression for LLM Attention
by: Yankun, Hong, et al.
Published: (2025)
by: Yankun, Hong, et al.
Published: (2025)
EvidenceMap: Learning Evidence Analysis to Unleash the Power of Small Language Models for Biomedical Question Answering
by: Zong, Chang, et al.
Published: (2025)
by: Zong, Chang, et al.
Published: (2025)
Similar Items
-
Retrieval-Based Multi-Label Legal Annotation: Extensible, Data-Efficient and Hallucination-Free
by: Zhang, Li, et al.
Published: (2026) -
Do LLMs Truly Understand When a Precedent Is Overruled?
by: Zhang, Li, et al.
Published: (2025) -
Mitigating Manipulation and Enhancing Persuasion: A Reflective Multi-Agent Approach for Legal Argument Generation
by: Zhang, Li, et al.
Published: (2025) -
Thinking Longer, Not Always Smarter: Evaluating LLM Capabilities in Hierarchical Legal Reasoning
by: Zhang, Li, et al.
Published: (2025) -
Using LLMs to Discover Legal Factors
by: Gray, Morgan, et al.
Published: (2024)