Generation Space Size: Understanding and Calibrating Open-Endedness of LLM Generations

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yu, Sunny, Jabbar, Ahmad, Hawkins, Robert, Jurafsky, Dan, Cheng, Myra
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914092303777792
author Yu, Sunny
Jabbar, Ahmad
Hawkins, Robert
Jurafsky, Dan
Cheng, Myra
author_facet Yu, Sunny
Jabbar, Ahmad
Hawkins, Robert
Jurafsky, Dan
Cheng, Myra
contents Different open-ended generation tasks require different degrees of output diversity. However, current LLMs are often miscalibrated. They collapse to overly homogeneous outputs for creative tasks and hallucinate diverse but incorrect responses for factual tasks. We argue that these two failure modes are unified by, and can both be addressed by, the notion of effective generation space size (GSS) -- the set of semantically distinct outputs a model considers for a prompt. We present GSSBench, a task suite of prompt pairs with ground-truth GSS relationships to assess different metrics and understand where models diverge from desired behavior. We find that hallucination detection metrics, particularly EigenScore, consistently outperform standard diversity and uncertainty quantification metrics, while using only model internals, providing interpretable insights into a model's internal task representations. We demonstrate three applications of GSS: (1) detecting prompt ambiguity and predicting clarification questions for better grounding, (2) interpreting overthinking and underthinking in reasoning models, and (3) steering models to expand their generation space to yield high-quality and diverse outputs.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12699
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generation Space Size: Understanding and Calibrating Open-Endedness of LLM Generations
Yu, Sunny
Jabbar, Ahmad
Hawkins, Robert
Jurafsky, Dan
Cheng, Myra
Computation and Language
Artificial Intelligence
Different open-ended generation tasks require different degrees of output diversity. However, current LLMs are often miscalibrated. They collapse to overly homogeneous outputs for creative tasks and hallucinate diverse but incorrect responses for factual tasks. We argue that these two failure modes are unified by, and can both be addressed by, the notion of effective generation space size (GSS) -- the set of semantically distinct outputs a model considers for a prompt. We present GSSBench, a task suite of prompt pairs with ground-truth GSS relationships to assess different metrics and understand where models diverge from desired behavior. We find that hallucination detection metrics, particularly EigenScore, consistently outperform standard diversity and uncertainty quantification metrics, while using only model internals, providing interpretable insights into a model's internal task representations. We demonstrate three applications of GSS: (1) detecting prompt ambiguity and predicting clarification questions for better grounding, (2) interpreting overthinking and underthinking in reasoning models, and (3) steering models to expand their generation space to yield high-quality and diverse outputs.
title Generation Space Size: Understanding and Calibrating Open-Endedness of LLM Generations
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.12699