Speaker Fuzzy Fingerprints: Benchmarking Text-Based Identification in Multiparty Dialogues

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ribeiro, Rui, Coheur, Luísa, Carvalho, Joao P.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912338358042624
author Ribeiro, Rui
Coheur, Luísa
Carvalho, Joao P.
author_facet Ribeiro, Rui
Coheur, Luísa
Carvalho, Joao P.
contents Speaker identification using voice recordings leverages unique acoustic features, but this approach fails when only textual data is available. Few approaches have attempted to tackle the problem of identifying speakers solely from text, and the existing ones have primarily relied on traditional methods. In this work, we explore the use of fuzzy fingerprints from large pre-trained models to improve text-based speaker identification. We integrate speaker-specific tokens and context-aware modeling, demonstrating that conversational context significantly boosts accuracy, reaching 70.6% on the Friends dataset and 67.7% on the Big Bang Theory dataset. Additionally, we show that fuzzy fingerprints can approximate full fine-tuning performance with fewer hidden units, offering improved interpretability. Finally, we analyze ambiguous utterances and propose a mechanism to detect speaker-agnostic lines. Our findings highlight key challenges and provide insights for future improvements in text-based speaker identification.
format Preprint
id arxiv_https___arxiv_org_abs_2504_14963
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Speaker Fuzzy Fingerprints: Benchmarking Text-Based Identification in Multiparty Dialogues
Ribeiro, Rui
Coheur, Luísa
Carvalho, Joao P.
Computation and Language
Artificial Intelligence
Computational Engineering, Finance, and Science
Machine Learning
Neural and Evolutionary Computing
Speaker identification using voice recordings leverages unique acoustic features, but this approach fails when only textual data is available. Few approaches have attempted to tackle the problem of identifying speakers solely from text, and the existing ones have primarily relied on traditional methods. In this work, we explore the use of fuzzy fingerprints from large pre-trained models to improve text-based speaker identification. We integrate speaker-specific tokens and context-aware modeling, demonstrating that conversational context significantly boosts accuracy, reaching 70.6% on the Friends dataset and 67.7% on the Big Bang Theory dataset. Additionally, we show that fuzzy fingerprints can approximate full fine-tuning performance with fewer hidden units, offering improved interpretability. Finally, we analyze ambiguous utterances and propose a mechanism to detect speaker-agnostic lines. Our findings highlight key challenges and provide insights for future improvements in text-based speaker identification.
title Speaker Fuzzy Fingerprints: Benchmarking Text-Based Identification in Multiparty Dialogues
topic Computation and Language
Artificial Intelligence
Computational Engineering, Finance, and Science
Machine Learning
Neural and Evolutionary Computing
url https://arxiv.org/abs/2504.14963