Evolutionary Search for Automated Design of Uncertainty Quantification Methods
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917383806910464 |
|---|---|
| author | Seleznyov, Mikhail Korbut, Daniil Moskvoretskii, Viktor Somov, Oleg Panchenko, Alexander Tutubalina, Elena |
| author_facet | Seleznyov, Mikhail Korbut, Daniil Moskvoretskii, Viktor Somov, Oleg Panchenko, Alexander Tutubalina, Elena |
| contents | Uncertainty quantification (UQ) methods for large language models are predominantly designed by hand based on domain knowledge and heuristics, limiting their scalability and generality. We apply LLM-powered evolutionary search to automatically discover unsupervised UQ methods represented as Python programs. On the task of atomic claim verification, our evolved methods outperform strong manually-designed baselines, achieving up to 6.7% relative ROC-AUC improvement across 9 datasets while generalizing robustly out-of-distribution. Qualitative analysis reveals that different LLMs employ qualitatively distinct evolutionary strategies: Claude models consistently design high-feature-count linear estimators, while Gpt-oss-120B gravitates toward simpler and more interpretable positional weighting schemes. Surprisingly, only Sonnet 4.5 and Opus 4.5 reliably leverage increased method complexity to improve performance -- Opus 4.6 shows an unexpected regression relative to its predecessor. Overall, our results indicate that LLM-powered evolutionary search is a promising paradigm for automated, interpretable hallucination detector design. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_03473 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Evolutionary Search for Automated Design of Uncertainty Quantification Methods Seleznyov, Mikhail Korbut, Daniil Moskvoretskii, Viktor Somov, Oleg Panchenko, Alexander Tutubalina, Elena Computation and Language Artificial Intelligence Uncertainty quantification (UQ) methods for large language models are predominantly designed by hand based on domain knowledge and heuristics, limiting their scalability and generality. We apply LLM-powered evolutionary search to automatically discover unsupervised UQ methods represented as Python programs. On the task of atomic claim verification, our evolved methods outperform strong manually-designed baselines, achieving up to 6.7% relative ROC-AUC improvement across 9 datasets while generalizing robustly out-of-distribution. Qualitative analysis reveals that different LLMs employ qualitatively distinct evolutionary strategies: Claude models consistently design high-feature-count linear estimators, while Gpt-oss-120B gravitates toward simpler and more interpretable positional weighting schemes. Surprisingly, only Sonnet 4.5 and Opus 4.5 reliably leverage increased method complexity to improve performance -- Opus 4.6 shows an unexpected regression relative to its predecessor. Overall, our results indicate that LLM-powered evolutionary search is a promising paradigm for automated, interpretable hallucination detector design. |
| title | Evolutionary Search for Automated Design of Uncertainty Quantification Methods |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2604.03473 |