An Empirical Study on Context Length for Open-Domain Dialog Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Xinyi, Lin, Zuoquan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913487900377088
author Shen, Xinyi
Lin, Zuoquan
author_facet Shen, Xinyi
Lin, Zuoquan
contents Transformer-based open-domain dialog models have become increasingly popular in recent years. These models typically represent context as a concatenation of a dialog history. However, there is no criterion to decide how many utterances should be kept adequate in a context. We try to figure out how the choice of context length affects the model. We experiment on three questions from coarse to fine: (i) Does longer context help model training? (ii) Is it necessary to change the training context length when dealing with dialogs of different context lengths? (iii) Do different dialog samples have the same preference for context length? Our experimental results show that context length, an often overlooked setting, deserves attention when implementing Transformer-based dialog models.
format Preprint
id arxiv_https___arxiv_org_abs_2409_00315
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle An Empirical Study on Context Length for Open-Domain Dialog Generation
Shen, Xinyi
Lin, Zuoquan
Computation and Language
Artificial Intelligence
Machine Learning
Transformer-based open-domain dialog models have become increasingly popular in recent years. These models typically represent context as a concatenation of a dialog history. However, there is no criterion to decide how many utterances should be kept adequate in a context. We try to figure out how the choice of context length affects the model. We experiment on three questions from coarse to fine: (i) Does longer context help model training? (ii) Is it necessary to change the training context length when dealing with dialogs of different context lengths? (iii) Do different dialog samples have the same preference for context length? Our experimental results show that context length, an often overlooked setting, deserves attention when implementing Transformer-based dialog models.
title An Empirical Study on Context Length for Open-Domain Dialog Generation
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2409.00315