No Single Best Model for Diversity: Learning a Router for Sample Diversity

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Liu, Yuhan, Xu, Fangyuan, Padmakumar, Vishakh, Ippolito, Daphne, Choi, Eunsol
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908950218145792
author Liu, Yuhan
Xu, Fangyuan
Padmakumar, Vishakh
Ippolito, Daphne
Choi, Eunsol
author_facet Liu, Yuhan
Xu, Fangyuan
Padmakumar, Vishakh
Ippolito, Daphne
Choi, Eunsol
contents When posed with prompts that permit a large number of valid answers, comprehensively generating them is the first step towards satisfying a wide range of users. In this paper, we study methods to elicit a comprehensive set of valid responses. To evaluate this, we introduce \textbf{diversity coverage}, a metric that measures the total quality scores assigned to each \textbf{unique} answer in the predicted answer set relative to the best possible answer set with the same number of answers. Using this metric, we evaluate 18 LLMs, finding no single model dominates at generating diverse responses to a wide range of open-ended prompts. Yet, per each prompt, there exists a model that outperforms all other models significantly at generating a diverse answer set. Motivated by this finding, we introduce a router that predicts the best model for each query. On NB-Wildchat, our trained router outperforms the single best model baseline (26.3% vs $23.8%). We further show generalization to an out-of-domain dataset (NB-Curated) as well as different answer-generation prompting strategies. Our work lays foundation for studying generating comprehensive answers when we have access to a suite of models.
format Preprint
id arxiv_https___arxiv_org_abs_2604_02319
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle No Single Best Model for Diversity: Learning a Router for Sample Diversity
Liu, Yuhan
Xu, Fangyuan
Padmakumar, Vishakh
Ippolito, Daphne
Choi, Eunsol
Computation and Language
When posed with prompts that permit a large number of valid answers, comprehensively generating them is the first step towards satisfying a wide range of users. In this paper, we study methods to elicit a comprehensive set of valid responses. To evaluate this, we introduce \textbf{diversity coverage}, a metric that measures the total quality scores assigned to each \textbf{unique} answer in the predicted answer set relative to the best possible answer set with the same number of answers. Using this metric, we evaluate 18 LLMs, finding no single model dominates at generating diverse responses to a wide range of open-ended prompts. Yet, per each prompt, there exists a model that outperforms all other models significantly at generating a diverse answer set. Motivated by this finding, we introduce a router that predicts the best model for each query. On NB-Wildchat, our trained router outperforms the single best model baseline (26.3% vs $23.8%). We further show generalization to an out-of-domain dataset (NB-Curated) as well as different answer-generation prompting strategies. Our work lays foundation for studying generating comprehensive answers when we have access to a suite of models.
title No Single Best Model for Diversity: Learning a Router for Sample Diversity
topic Computation and Language
url https://arxiv.org/abs/2604.02319