Supporting Humans in Evaluating AI Summaries of Legal Depositions
Fuente:
arXiv
Saved in:
| Main Authors: | Farzi, Naghmeh, Dietz, Laura, Lewis, Dave D. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Best in Tau@LLMJudge: Criteria-Based Relevance Evaluation with Llama3
by: Farzi, Naghmeh, et al.
Published: (2024)
by: Farzi, Naghmeh, et al.
Published: (2024)
Does UMBRELA Work on Other LLMs?
by: Farzi, Naghmeh, et al.
Published: (2025)
by: Farzi, Naghmeh, et al.
Published: (2025)
Criteria-Based LLM Relevance Judgments
by: Farzi, Naghmeh, et al.
Published: (2025)
by: Farzi, Naghmeh, et al.
Published: (2025)
An Exam-based Evaluation Approach Beyond Traditional Relevance Judgments
by: Farzi, Naghmeh, et al.
Published: (2024)
by: Farzi, Naghmeh, et al.
Published: (2024)
Diagnosing LLM-based Rerankers in Cold-Start Recommender Systems: Coverage, Exposure and Practical Mitigations
by: Lemdiasova, Ekaterina, et al.
Published: (2026)
by: Lemdiasova, Ekaterina, et al.
Published: (2026)
LLM-based relevance assessment still can't replace human relevance assessment
by: Clarke, Charles L. A., et al.
Published: (2024)
by: Clarke, Charles L. A., et al.
Published: (2024)
LLM-Evaluation Tropes: Perspectives on the Validity of LLM-Evaluations
by: Dietz, Laura, et al.
Published: (2025)
by: Dietz, Laura, et al.
Published: (2025)
AI for Statutory Simplification: A Comprehensive State Legal Corpus and Labor Benchmark
by: Hariri, Emaan, et al.
Published: (2025)
by: Hariri, Emaan, et al.
Published: (2025)
Agentic Search in the Wild: Intents and Trajectory Dynamics from 14M+ Real Search Requests
by: Ning, Jingjie, et al.
Published: (2026)
by: Ning, Jingjie, et al.
Published: (2026)
AmharicIR+Instr: A Two-Dataset Resource for Neural Retrieval and Instruction Tuning
by: Yeshambel, Tilahun, et al.
Published: (2026)
by: Yeshambel, Tilahun, et al.
Published: (2026)
Association via Entropy Reduction
by: Gamst, Anthony, et al.
Published: (2025)
by: Gamst, Anthony, et al.
Published: (2025)
SIGIR 2025 -- LiveRAG Challenge Report
by: Carmel, David, et al.
Published: (2025)
by: Carmel, David, et al.
Published: (2025)
Utilizing Large Language Models for Named Entity Recognition in Traditional Chinese Medicine against COVID-19 Literature: Comparative Study
by: Tong, Xu, et al.
Published: (2024)
by: Tong, Xu, et al.
Published: (2024)
UNIQORN: Unified Question Answering over RDF Knowledge Graphs and Natural Language Text
by: Pramanik, Soumajit, et al.
Published: (2021)
by: Pramanik, Soumajit, et al.
Published: (2021)
Rag Performance Prediction for Question Answering
by: Dado, Or, et al.
Published: (2026)
by: Dado, Or, et al.
Published: (2026)
From BM25 to Corrective RAG: Benchmarking Retrieval Strategies for Text-and-Table Documents
by: Akarsu, Meftun, et al.
Published: (2026)
by: Akarsu, Meftun, et al.
Published: (2026)
PyTerrier-GenRank: The PyTerrier Plugin for Reranking with Large Language Models
by: Dhole, Kaustubh D.
Published: (2024)
by: Dhole, Kaustubh D.
Published: (2024)
Beyond Negation Detection: Comprehensive Assertion Detection Models for Clinical NLP
by: Kocaman, Veysel, et al.
Published: (2025)
by: Kocaman, Veysel, et al.
Published: (2025)
SCTc-TE: A Comprehensive Formulation and Benchmark for Temporal Event Forecasting
by: Ma, Yunshan, et al.
Published: (2023)
by: Ma, Yunshan, et al.
Published: (2023)
Insider Knowledge: How Much Can RAG Systems Gain from Evaluation Secrets?
by: Dietz, Laura, et al.
Published: (2026)
by: Dietz, Laura, et al.
Published: (2026)
Towards Reliable Retrieval in RAG Systems for Large Legal Datasets
by: Reuter, Markus, et al.
Published: (2025)
by: Reuter, Markus, et al.
Published: (2025)
All for law and law for all: Adaptive RAG Pipeline for Legal Research
by: Keisha, Figarri, et al.
Published: (2025)
by: Keisha, Figarri, et al.
Published: (2025)
Falkor-IRAC: Graph-Constrained Generation for Verified Legal Reasoning in Indian Judicial AI
by: Bose, Joy
Published: (2026)
by: Bose, Joy
Published: (2026)
Reproduction Beyond Benchmarks: ConstBERT and ColBERT-v2 Across Backends and Query Distributions
by: Ghosh, Utshab Kumar, et al.
Published: (2026)
by: Ghosh, Utshab Kumar, et al.
Published: (2026)
Thinking Broad, Acting Fast: Latent Reasoning Distillation from Multi-Perspective Chain-of-Thought for E-Commerce Relevance
by: Qiu, Baopu, et al.
Published: (2026)
by: Qiu, Baopu, et al.
Published: (2026)
Association Is Not Similarity: Learning Corpus-Specific Associations for Multi-Hop Retrieval
by: Dury, Jason
Published: (2026)
by: Dury, Jason
Published: (2026)
A Systematic Review of Generative AI for Teaching and Learning Practice
by: Ogunleye, Bayode, et al.
Published: (2024)
by: Ogunleye, Bayode, et al.
Published: (2024)
Cross-Domain Keyword Extraction with Keyness Patterns
by: Zhou, Dongmei, et al.
Published: (2024)
by: Zhou, Dongmei, et al.
Published: (2024)
Generative Query Reformulation Using Ensemble Prompting, Document Fusion, and Relevance Feedback
by: Dhole, Kaustubh D., et al.
Published: (2024)
by: Dhole, Kaustubh D., et al.
Published: (2024)
Iterative NLP Query Refinement for Enhancing Domain-Specific Information Retrieval: A Case Study in Career Services
by: Peimani, Elham, et al.
Published: (2024)
by: Peimani, Elham, et al.
Published: (2024)
Legal RAG Bench: an end-to-end benchmark for legal RAG
by: Butler, Abdur-Rahman, et al.
Published: (2026)
by: Butler, Abdur-Rahman, et al.
Published: (2026)
TeroSeek: An AI-Powered Knowledge Base and Retrieval Generation Platform for Terpenoid Research
by: Kang, Xu, et al.
Published: (2025)
by: Kang, Xu, et al.
Published: (2025)
MLDocRAG: Multimodal Long-Context Document Retrieval Augmented Generation
by: Zhang, Yongyue, et al.
Published: (2026)
by: Zhang, Yongyue, et al.
Published: (2026)
Temporal Decay of Co-Citation Predictability: A 20-Year Statute Retrieval Benchmark from 396M Ukrainian Court Citations
by: Ovcharov, Volodymyr
Published: (2026)
by: Ovcharov, Volodymyr
Published: (2026)
Temporal Leakage in Search-Engine Date-Filtered Web Retrieval: A Retrospective Forecasting Case Study
by: Lahib, Ali El, et al.
Published: (2026)
by: Lahib, Ali El, et al.
Published: (2026)
A Chain-of-Thought Approach to Semantic Query Categorization in e-Commerce Taxonomies
by: Duraj, Jetlir, et al.
Published: (2026)
by: Duraj, Jetlir, et al.
Published: (2026)
VIRAASAT: Traversing Novel Paths for Indian Cultural Reasoning
by: Surana, Harshul Raj, et al.
Published: (2026)
by: Surana, Harshul Raj, et al.
Published: (2026)
The Reasoning Bottleneck in Graph-RAG: Structured Prompting and Context Compression for Multi-Hop QA
by: Zarrinkia, Yasaman, et al.
Published: (2026)
by: Zarrinkia, Yasaman, et al.
Published: (2026)
MAB-DQA: Addressing Query Aspect Importance in Document Question Answering with Multi-Armed Bandits
by: Xiang, Yixin, et al.
Published: (2026)
by: Xiang, Yixin, et al.
Published: (2026)
Augmented Relevance Datasets with Fine-Tuned Small LLMs
by: Fitte-Rey, Quentin, et al.
Published: (2025)
by: Fitte-Rey, Quentin, et al.
Published: (2025)
Similar Items
-
Best in Tau@LLMJudge: Criteria-Based Relevance Evaluation with Llama3
by: Farzi, Naghmeh, et al.
Published: (2024) -
Does UMBRELA Work on Other LLMs?
by: Farzi, Naghmeh, et al.
Published: (2025) -
Criteria-Based LLM Relevance Judgments
by: Farzi, Naghmeh, et al.
Published: (2025) -
An Exam-based Evaluation Approach Beyond Traditional Relevance Judgments
by: Farzi, Naghmeh, et al.
Published: (2024) -
Diagnosing LLM-based Rerankers in Cold-Start Recommender Systems: Coverage, Exposure and Practical Mitigations
by: Lemdiasova, Ekaterina, et al.
Published: (2026)