STEER: Inference-Time Risk Control via Constrained Quality-Diversity Search

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yang, Eric, Lee, Jong Ha, Amar, Jonathan, Ye, Elissa, Jia, Yugang
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918320614146048
author Yang, Eric
Lee, Jong Ha
Amar, Jonathan
Ye, Elissa
Jia, Yugang
author_facet Yang, Eric
Lee, Jong Ha
Amar, Jonathan
Ye, Elissa
Jia, Yugang
contents Large Language Models (LLMs) trained for average correctness often exhibit mode collapse, producing narrow decision behaviors on tasks where multiple responses may be reasonable. This limitation is particularly problematic in ordinal decision settings such as clinical triage, where standard alignment removes the ability to trade off specificity and sensitivity (the ROC operating point) based on contextual constraints. We propose STEER (Steerable Tuning via Evolutionary Ensemble Refinement), a training-free framework that reintroduces this tunable control. STEER constructs a population of natural-language personas through an offline, constrained quality-diversity search that promotes behavioral coverage while enforcing minimum safety, reasoning, and stability thresholds. At inference time, STEER exposes a single, interpretable control parameter that maps a user-specified risk percentile to a selected persona, yielding a monotonic adjustment of decision conservativeness. On two clinical triage benchmarks, STEER achieves broader behavioral coverage compared to temperature-based sampling and static persona ensembles. Compared to a representative post-training method, STEER maintains substantially higher accuracy on unambiguous urgent cases while providing comparable control over ambiguous decisions. These results demonstrate STEER as a safety-preserving paradigm for risk control, capable of steering behavior without compromising domain competence.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02862
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle STEER: Inference-Time Risk Control via Constrained Quality-Diversity Search
Yang, Eric
Lee, Jong Ha
Amar, Jonathan
Ye, Elissa
Jia, Yugang
Artificial Intelligence
Machine Learning
Large Language Models (LLMs) trained for average correctness often exhibit mode collapse, producing narrow decision behaviors on tasks where multiple responses may be reasonable. This limitation is particularly problematic in ordinal decision settings such as clinical triage, where standard alignment removes the ability to trade off specificity and sensitivity (the ROC operating point) based on contextual constraints. We propose STEER (Steerable Tuning via Evolutionary Ensemble Refinement), a training-free framework that reintroduces this tunable control. STEER constructs a population of natural-language personas through an offline, constrained quality-diversity search that promotes behavioral coverage while enforcing minimum safety, reasoning, and stability thresholds. At inference time, STEER exposes a single, interpretable control parameter that maps a user-specified risk percentile to a selected persona, yielding a monotonic adjustment of decision conservativeness. On two clinical triage benchmarks, STEER achieves broader behavioral coverage compared to temperature-based sampling and static persona ensembles. Compared to a representative post-training method, STEER maintains substantially higher accuracy on unambiguous urgent cases while providing comparable control over ambiguous decisions. These results demonstrate STEER as a safety-preserving paradigm for risk control, capable of steering behavior without compromising domain competence.
title STEER: Inference-Time Risk Control via Constrained Quality-Diversity Search
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2602.02862