Saved in:
Bibliographic Details
Main Authors: Zhou, Yuxuan, Keuper, Margret, Fritz, Mario
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2408.13586
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915094189834240
author Zhou, Yuxuan
Keuper, Margret
Fritz, Mario
author_facet Zhou, Yuxuan
Keuper, Margret
Fritz, Mario
contents Sampling-based decoding strategies have been widely adopted for Large Language Models (LLMs) in numerous applications, targeting a balance between diversity and quality via temperature tuning and tail truncation. Considering the strong dependency of the candidate next tokens on different prefixes, recent studies propose to adaptively truncate the tail of LLMs' predicted distribution. Although improved results have been reported with these methods on open-ended text generation tasks, the results are highly dependent on the curated parameters and the limited exemplar text. In this paper, we propose a systematic way to estimate the capacity of a truncation sampling method by considering the trade-off between diversity and risk at each decoding step, based on our collected prefix tree which preserves the context of a full sentence. Our work offers a comprehensive comparison of existing truncation sampling methods and serves as a practical user guideline for their parameter selection.
format Preprint
id arxiv_https___arxiv_org_abs_2408_13586
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Balancing Diversity and Risk in LLM Sampling: How to Select Your Method and Parameter for Open-Ended Text Generation
Zhou, Yuxuan
Keuper, Margret
Fritz, Mario
Computation and Language
Artificial Intelligence
Sampling-based decoding strategies have been widely adopted for Large Language Models (LLMs) in numerous applications, targeting a balance between diversity and quality via temperature tuning and tail truncation. Considering the strong dependency of the candidate next tokens on different prefixes, recent studies propose to adaptively truncate the tail of LLMs' predicted distribution. Although improved results have been reported with these methods on open-ended text generation tasks, the results are highly dependent on the curated parameters and the limited exemplar text. In this paper, we propose a systematic way to estimate the capacity of a truncation sampling method by considering the trade-off between diversity and risk at each decoding step, based on our collected prefix tree which preserves the context of a full sentence. Our work offers a comprehensive comparison of existing truncation sampling methods and serves as a practical user guideline for their parameter selection.
title Balancing Diversity and Risk in LLM Sampling: How to Select Your Method and Parameter for Open-Ended Text Generation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2408.13586