When Gradients Collide: Failure Modes of Multi-Objective Prompt Optimization for LLM Judges
Fuente:
arXiv
Saved in:
| Main Authors: | , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911716734926848 |
|---|---|
| author | Darshan, Parth Divekar, Abhishek |
| author_facet | Darshan, Parth Divekar, Abhishek |
| contents | Customizing an LLM judge to a specific task or domain often involves optimizing its prompt across multiple evaluation criteria simultaneously. Textual gradient methods automate this for a single judge criterion, however they produce natural-language critiques, not numerical vectors. Thus, the conflict-resolution toolkit of multi-task learning (PCGrad, MGDA) doesn't apply to the multi-objective textual gradient setting. We test five decomposition modes of textual gradient optimizers by varying how much cross-task information the loss, gradient and optimizer LLMs share. In 6 of 10 configurations, we observe that optimization never improves over the initial prompt. Gradient specificity drops by 59% (from 9.0 to 3.7) when the gradient LLM processes multiple criteria jointly. Separately, we observe that naively combining per-task instructions into a single prompt degrades Spearman's rho by -5.3%. These results identify two separable failure modes: optimization-time gradient dilution and inference-time instruction interference, which together constrain the design space for multi-objective judge customization using textual feedback. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_26046 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | When Gradients Collide: Failure Modes of Multi-Objective Prompt Optimization for LLM Judges Darshan, Parth Divekar, Abhishek Computation and Language Artificial Intelligence Machine Learning Multiagent Systems Software Engineering I.2.7; I.2.6; I.2.4; I.2.8 Customizing an LLM judge to a specific task or domain often involves optimizing its prompt across multiple evaluation criteria simultaneously. Textual gradient methods automate this for a single judge criterion, however they produce natural-language critiques, not numerical vectors. Thus, the conflict-resolution toolkit of multi-task learning (PCGrad, MGDA) doesn't apply to the multi-objective textual gradient setting. We test five decomposition modes of textual gradient optimizers by varying how much cross-task information the loss, gradient and optimizer LLMs share. In 6 of 10 configurations, we observe that optimization never improves over the initial prompt. Gradient specificity drops by 59% (from 9.0 to 3.7) when the gradient LLM processes multiple criteria jointly. Separately, we observe that naively combining per-task instructions into a single prompt degrades Spearman's rho by -5.3%. These results identify two separable failure modes: optimization-time gradient dilution and inference-time instruction interference, which together constrain the design space for multi-objective judge customization using textual feedback. |
| title | When Gradients Collide: Failure Modes of Multi-Objective Prompt Optimization for LLM Judges |
| topic | Computation and Language Artificial Intelligence Machine Learning Multiagent Systems Software Engineering I.2.7; I.2.6; I.2.4; I.2.8 |
| url | https://arxiv.org/abs/2605.26046 |