Exploiting LLM-as-a-Judge Disposition on Free Text Legal QA via Prompt Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Elganayni, Mohamed Hesham, Chen, Runsheng, Nagl, Sebastian, Grabmair, Matthias |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Thinking Longer, Not Always Smarter: Evaluating LLM Capabilities in Hierarchical Legal Reasoning
by: Zhang, Li, et al.
Published: (2025)
by: Zhang, Li, et al.
Published: (2025)
The Judge Variable: Challenging Judge-Agnostic Legal Judgment Prediction
by: Zambrano, Guillaume
Published: (2025)
by: Zambrano, Guillaume
Published: (2025)
BLT: Can Large Language Models Handle Basic Legal Text?
by: Blair-Stanek, Andrew, et al.
Published: (2023)
by: Blair-Stanek, Andrew, et al.
Published: (2023)
Evaluating LLM Prompts for Data Augmentation in Multi-label Classification of Ecological Texts
by: Glazkova, Anna, et al.
Published: (2024)
by: Glazkova, Anna, et al.
Published: (2024)
Prompting from the bench: Large-scale pretraining is not sufficient to prepare LLMs for ordinary meaning analysis
by: Purushothama, Abhishek, et al.
Published: (2025)
by: Purushothama, Abhishek, et al.
Published: (2025)
BenCSSmark: Making the Social Sciences Count in LLM Research
by: Chatelain, Arnault, et al.
Published: (2026)
by: Chatelain, Arnault, et al.
Published: (2026)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
by: Ashuach, Tomer, et al.
Published: (2025)
by: Ashuach, Tomer, et al.
Published: (2025)
By Their Fruits You Will Know Them: Comparing Formalizations of Law by the Decisions They Encode
by: Vernie, Julius, et al.
Published: (2026)
by: Vernie, Julius, et al.
Published: (2026)
Improving QA Model Performance with Cartographic Inoculation
by: Chen, Allen, et al.
Published: (2024)
by: Chen, Allen, et al.
Published: (2024)
Reshaping Free-Text Radiology Notes Into Structured Reports With Generative Transformers
by: Bergomi, Laura, et al.
Published: (2024)
by: Bergomi, Laura, et al.
Published: (2024)
How Much Does Persuasion Strategy Matter? LLM-Annotated Evidence from Charitable Donation Dialogues
by: Petrova, Tatiana, et al.
Published: (2026)
by: Petrova, Tatiana, et al.
Published: (2026)
Can Large Language Models Imitate Human Speech for Clinical Assessment? LLM-Driven Data Augmentation for Cognitive Score Prediction
by: Ketir, Si-Belkacem Yamine, et al.
Published: (2026)
by: Ketir, Si-Belkacem Yamine, et al.
Published: (2026)
Unified Game Moderation: Soft-Prompting and LLM-Assisted Label Transfer for Resource-Efficient Toxicity Detection
by: Yang, Zachary, et al.
Published: (2025)
by: Yang, Zachary, et al.
Published: (2025)
MAWARITH: A Dataset and Benchmark for Legal Inheritance Reasoning with LLMs
by: Bouchekif, Abdessalam, et al.
Published: (2026)
by: Bouchekif, Abdessalam, et al.
Published: (2026)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025)
by: Peters, Sydney, et al.
Published: (2025)
UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning
by: Ovcharov, Volodymyr
Published: (2026)
by: Ovcharov, Volodymyr
Published: (2026)
Decoding-Free Sampling Strategies for LLM Marginalization
by: Pohl, David, et al.
Published: (2025)
by: Pohl, David, et al.
Published: (2025)
Multi-Method Validation of Large Language Model Medical Translation Across High- and Low-Resource Languages
by: Anyaegbuna, Chukwuebuka, et al.
Published: (2026)
by: Anyaegbuna, Chukwuebuka, et al.
Published: (2026)
Survey and Experiments on Mental Disorder Detection via Social Media: From Large Language Models and RAG to Agents
by: Ge, Zhuohan, et al.
Published: (2025)
by: Ge, Zhuohan, et al.
Published: (2025)
Comparing Complex Concepts with Transformers: Matching Patent Claims Against Natural Language Text
by: Blume, Matthias, et al.
Published: (2024)
by: Blume, Matthias, et al.
Published: (2024)
The Foundational Capabilities of Large Language Models in Predicting Postoperative Risks Using Clinical Notes
by: Alba, Charles, et al.
Published: (2024)
by: Alba, Charles, et al.
Published: (2024)
What distinguishes conspiracy from critical narratives? A computational analysis of oppositional discourse
by: Korenčić, Damir, et al.
Published: (2024)
by: Korenčić, Damir, et al.
Published: (2024)
Using Letter Positional Probabilities to Assess Word Complexity
by: Dalvean, Michael
Published: (2024)
by: Dalvean, Michael
Published: (2024)
The Provenance Gap in Clinical AI: Evidence-Traceable Temporal Knowledge Graphs for Rare Disease Reasoning
by: Ahmed, Md Shamim, et al.
Published: (2026)
by: Ahmed, Md Shamim, et al.
Published: (2026)
LLMs Generate Kitsch
by: Klinge, Xenia, et al.
Published: (2026)
by: Klinge, Xenia, et al.
Published: (2026)
Integrating clinical reasoning into large language model-based diagnosis through etiology-aware attention steering
by: Li, Peixian, et al.
Published: (2025)
by: Li, Peixian, et al.
Published: (2025)
The Representational Alignment between Humans and Language Models is implicitly driven by a Concreteness Effect
by: Iaia, Cosimo, et al.
Published: (2025)
by: Iaia, Cosimo, et al.
Published: (2025)
Cheap Learning: Maximising Performance of Language Models for Social Data Science Using Minimal Data
by: Castro-Gonzalez, Leonardo, et al.
Published: (2024)
by: Castro-Gonzalez, Leonardo, et al.
Published: (2024)
Comparative Study of Large Language Models on Chinese Film Script Continuation: An Empirical Analysis Based on GPT-5.2 and Qwen-Max
by: Cao, Yuxuan, et al.
Published: (2026)
by: Cao, Yuxuan, et al.
Published: (2026)
Forgotten Words: Benchmarking NeoBERT for Dementia Detection in Low-Resource Conversational Filipino and English Speech
by: Floresca, Rez Samantha Z., et al.
Published: (2026)
by: Floresca, Rez Samantha Z., et al.
Published: (2026)
Collective Memory and Narrative Cohesion: A Computational Study of Palestinian Refugee Oral Histories in Lebanon
by: Awwad, Ghadeer, et al.
Published: (2025)
by: Awwad, Ghadeer, et al.
Published: (2025)
EfficientQA : a RoBERTa Based Phrase-Indexed Question-Answering System
by: Chaybouti, Sofian, et al.
Published: (2021)
by: Chaybouti, Sofian, et al.
Published: (2021)
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
by: Chehbouni, Khaoula, et al.
Published: (2025)
by: Chehbouni, Khaoula, et al.
Published: (2025)
LLM-as-a-Judge: Rapid Evaluation of Legal Document Recommendation for Retrieval-Augmented Generation
by: Pradhan, Anu, et al.
Published: (2025)
by: Pradhan, Anu, et al.
Published: (2025)
APIO: Automatic Prompt Induction and Optimization for Grammatical Error Correction and Text Simplification
by: Chernodub, Artem, et al.
Published: (2025)
by: Chernodub, Artem, et al.
Published: (2025)
ARGUS: Seeing the Influence of Narrative Features on Persuasion in Argumentative Texts
by: Nabhani, Sara, et al.
Published: (2026)
by: Nabhani, Sara, et al.
Published: (2026)
HumanLLM: Benchmarking and Improving LLM Anthropomorphism via Human Cognitive Patterns
by: Wang, Xintao, et al.
Published: (2026)
by: Wang, Xintao, et al.
Published: (2026)
Model selection meets clinical semantics: Optimizing ICD-10-CM prediction via LLM-as-Judge evaluation, redundancy-aware sampling, and section-aware fine-tuning
by: Dai, Hong-Jie, et al.
Published: (2025)
by: Dai, Hong-Jie, et al.
Published: (2025)
Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text Perceptions
by: Orlikowski, Matthias, et al.
Published: (2025)
by: Orlikowski, Matthias, et al.
Published: (2025)
Citation Grounding: Detecting and Reducing LLM Citation Hallucinations via Legal Citation Graphs
by: Ovcharov, Volodymyr
Published: (2026)
by: Ovcharov, Volodymyr
Published: (2026)
Similar Items
-
Thinking Longer, Not Always Smarter: Evaluating LLM Capabilities in Hierarchical Legal Reasoning
by: Zhang, Li, et al.
Published: (2025) -
The Judge Variable: Challenging Judge-Agnostic Legal Judgment Prediction
by: Zambrano, Guillaume
Published: (2025) -
BLT: Can Large Language Models Handle Basic Legal Text?
by: Blair-Stanek, Andrew, et al.
Published: (2023) -
Evaluating LLM Prompts for Data Augmentation in Multi-label Classification of Ecological Texts
by: Glazkova, Anna, et al.
Published: (2024) -
Prompting from the bench: Large-scale pretraining is not sufficient to prepare LLMs for ordinary meaning analysis
by: Purushothama, Abhishek, et al.
Published: (2025)