Exploiting LLM-as-a-Judge Disposition on Free Text Legal QA via Prompt Optimization

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Elganayni, Mohamed Hesham, Chen, Runsheng, Nagl, Sebastian, Grabmair, Matthias
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913057662304256
author Elganayni, Mohamed Hesham
Chen, Runsheng
Nagl, Sebastian
Grabmair, Matthias
author_facet Elganayni, Mohamed Hesham
Chen, Runsheng
Nagl, Sebastian
Grabmair, Matthias
contents This work explores the role of prompt design and judge selection in LLM-as-a-Judge evaluations of free text legal question answering. We examine whether automatic task prompt optimization improves over human-centered design, whether optimization effectiveness varies by judge feedback style, and whether optimized prompts transfer across judges. We systematically address these questions on the LEXam benchmark by optimizing task prompts using the ProTeGi method with feedback from two judges (Qwen3-32B, DeepSeek-V3) across four task models, and then testing cross-judge transfer. Automatic optimization consistently outperforms the baseline, with lenient judge feedback yielding higher and more consistent gains than strict judge feedback. Prompts optimized with lenient feedback transfer better to strict judges than the reverse direction. Analysis reveals that lenient judges provide permissive feedback, yielding prompts with broader applicability, whereas strict judges produce restrictive feedback, leading to judge-specific overfitting. Our findings demonstrate algorithmically optimizing prompts on training data can outperform human-centered prompt design and that judges' dispositions during optimization shape prompt generalizability.
format Preprint
id arxiv_https___arxiv_org_abs_2604_20726
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Exploiting LLM-as-a-Judge Disposition on Free Text Legal QA via Prompt Optimization
Elganayni, Mohamed Hesham
Chen, Runsheng
Nagl, Sebastian
Grabmair, Matthias
Computation and Language
Artificial Intelligence
I.2.7; J.7
This work explores the role of prompt design and judge selection in LLM-as-a-Judge evaluations of free text legal question answering. We examine whether automatic task prompt optimization improves over human-centered design, whether optimization effectiveness varies by judge feedback style, and whether optimized prompts transfer across judges. We systematically address these questions on the LEXam benchmark by optimizing task prompts using the ProTeGi method with feedback from two judges (Qwen3-32B, DeepSeek-V3) across four task models, and then testing cross-judge transfer. Automatic optimization consistently outperforms the baseline, with lenient judge feedback yielding higher and more consistent gains than strict judge feedback. Prompts optimized with lenient feedback transfer better to strict judges than the reverse direction. Analysis reveals that lenient judges provide permissive feedback, yielding prompts with broader applicability, whereas strict judges produce restrictive feedback, leading to judge-specific overfitting. Our findings demonstrate algorithmically optimizing prompts on training data can outperform human-centered prompt design and that judges' dispositions during optimization shape prompt generalizability.
title Exploiting LLM-as-a-Judge Disposition on Free Text Legal QA via Prompt Optimization
topic Computation and Language
Artificial Intelligence
I.2.7; J.7
url https://arxiv.org/abs/2604.20726