Investigating the Representation of Backchannels and Fillers in Fine-tuned Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Yu, Lao, Leyi, Huang, Langchu, Skantze, Gabriel, Xu, Yang, Buschmeier, Hendrik
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911622144983040
author Wang, Yu
Lao, Leyi
Huang, Langchu
Skantze, Gabriel
Xu, Yang
Buschmeier, Hendrik
author_facet Wang, Yu
Lao, Leyi
Huang, Langchu
Skantze, Gabriel
Xu, Yang
Buschmeier, Hendrik
contents Backchannels and fillers are important linguistic expressions in dialogue, but often treated as 'noise' to be bypassed in modern transformer-based language models (LMs). Here, we study how they are represented in LMs using three fine-tuning strategies on three dialogue corpora in English and Japanese, in which backchannels and fillers are both preserved and annotated. This allows us to investigate how fine-tuning can help LMs learn these representations. We first apply clustering analysis to the learnt representation of backchannels and fillers, and find increased silhouette scores in representations from fine-tuned models, which suggests that fine-tuning enables LMs to distinguish the nuanced semantic variation in different backchannel and filler use. We also employ natural language generation metrics and qualitative analyses to verify that utterances produced by fine-tuned LMs resemble those produced by humans more closely. Our findings suggest the potential for transforming general LMs into conversational LMs that can produce human-like language more adequately.
format Preprint
id arxiv_https___arxiv_org_abs_2509_20237
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Investigating the Representation of Backchannels and Fillers in Fine-tuned Language Models
Wang, Yu
Lao, Leyi
Huang, Langchu
Skantze, Gabriel
Xu, Yang
Buschmeier, Hendrik
Computation and Language
Backchannels and fillers are important linguistic expressions in dialogue, but often treated as 'noise' to be bypassed in modern transformer-based language models (LMs). Here, we study how they are represented in LMs using three fine-tuning strategies on three dialogue corpora in English and Japanese, in which backchannels and fillers are both preserved and annotated. This allows us to investigate how fine-tuning can help LMs learn these representations. We first apply clustering analysis to the learnt representation of backchannels and fillers, and find increased silhouette scores in representations from fine-tuned models, which suggests that fine-tuning enables LMs to distinguish the nuanced semantic variation in different backchannel and filler use. We also employ natural language generation metrics and qualitative analyses to verify that utterances produced by fine-tuned LMs resemble those produced by humans more closely. Our findings suggest the potential for transforming general LMs into conversational LMs that can produce human-like language more adequately.
title Investigating the Representation of Backchannels and Fillers in Fine-tuned Language Models
topic Computation and Language
url https://arxiv.org/abs/2509.20237