Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jain, Shomik, Lanchantin, Jack, Nickel, Maximilian, Ross, Candace, Ullrich, Karen, Wilson, Ashia, Watson-Daniels, Jamelle
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910157178404864
author Jain, Shomik
Lanchantin, Jack
Nickel, Maximilian
Ross, Candace
Ullrich, Karen
Wilson, Ashia
Watson-Daniels, Jamelle
author_facet Jain, Shomik
Lanchantin, Jack
Nickel, Maximilian
Ross, Candace
Ullrich, Karen
Wilson, Ashia
Watson-Daniels, Jamelle
contents Large language models often generate homogeneous outputs, but whether this is problematic depends on the specific task. For objective math tasks, responses may vary in terms of problem-solving strategy but should maintain the same verifiable answer. Whereas, for creative writing tasks, we often expect variation in key narrative components (e.g. plot, setting, etc.) beyond mere vocabulary diversity. Prior work on homogenization rarely conceptualizes diversity in a task-dependent way. We address this gap with four contributions: (1) a task taxonomy with distinct notions of functional diversity -- whether a user would perceive two responses as meaningfully different for a given task; (2) a small user study validating that the taxonomy aligns with human perception of functional diversity; (3) a task-dependent sampling technique that increases diversity only where homogenization is undesired; (4) evidence challenging the perceived diversity-quality trade-off, showing it may stem from mis-conceptualizing both diversity and quality in a task-agnostic way.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21267
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework
Jain, Shomik
Lanchantin, Jack
Nickel, Maximilian
Ross, Candace
Ullrich, Karen
Wilson, Ashia
Watson-Daniels, Jamelle
Computation and Language
Computers and Society
Large language models often generate homogeneous outputs, but whether this is problematic depends on the specific task. For objective math tasks, responses may vary in terms of problem-solving strategy but should maintain the same verifiable answer. Whereas, for creative writing tasks, we often expect variation in key narrative components (e.g. plot, setting, etc.) beyond mere vocabulary diversity. Prior work on homogenization rarely conceptualizes diversity in a task-dependent way. We address this gap with four contributions: (1) a task taxonomy with distinct notions of functional diversity -- whether a user would perceive two responses as meaningfully different for a given task; (2) a small user study validating that the taxonomy aligns with human perception of functional diversity; (3) a task-dependent sampling technique that increases diversity only where homogenization is undesired; (4) evidence challenging the perceived diversity-quality trade-off, showing it may stem from mis-conceptualizing both diversity and quality in a task-agnostic way.
title Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework
topic Computation and Language
Computers and Society
url https://arxiv.org/abs/2509.21267