Saved in:
Bibliographic Details
Main Authors: Penzo, Nicolò, Sajedinia, Maryam, Lepri, Bruno, Tonelli, Sara, Guerini, Marco
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2409.18602
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917363746603008
author Penzo, Nicolò
Sajedinia, Maryam
Lepri, Bruno
Tonelli, Sara
Guerini, Marco
author_facet Penzo, Nicolò
Sajedinia, Maryam
Lepri, Bruno
Tonelli, Sara
Guerini, Marco
contents Assessing the performance of systems to classify Multi-Party Conversations (MPC) is challenging due to the interconnection between linguistic and structural characteristics of conversations. Conventional evaluation methods often overlook variances in model behavior across different levels of structural complexity on interaction graphs. In this work, we propose a methodological pipeline to investigate model performance across specific structural attributes of conversations. As a proof of concept we focus on Response Selection and Addressee Recognition tasks, to diagnose model weaknesses. To this end, we extract representative diagnostic subdatasets with a fixed number of users and a good structural variety from a large and open corpus of online MPCs. We further frame our work in terms of data minimization, avoiding the use of original usernames to preserve privacy, and propose alternatives to using original text messages. Results show that response selection relies more on the textual content of conversations, while addressee recognition requires capturing their structural dimension. Using an LLM in a zero-shot setting, we further highlight how sensitivity to prompt variations is task-dependent.
format Preprint
id arxiv_https___arxiv_org_abs_2409_18602
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Do LLMs suffer from Multi-Party Hangover? A Diagnostic Approach to Addressee Recognition and Response Selection in Conversations
Penzo, Nicolò
Sajedinia, Maryam
Lepri, Bruno
Tonelli, Sara
Guerini, Marco
Computation and Language
Assessing the performance of systems to classify Multi-Party Conversations (MPC) is challenging due to the interconnection between linguistic and structural characteristics of conversations. Conventional evaluation methods often overlook variances in model behavior across different levels of structural complexity on interaction graphs. In this work, we propose a methodological pipeline to investigate model performance across specific structural attributes of conversations. As a proof of concept we focus on Response Selection and Addressee Recognition tasks, to diagnose model weaknesses. To this end, we extract representative diagnostic subdatasets with a fixed number of users and a good structural variety from a large and open corpus of online MPCs. We further frame our work in terms of data minimization, avoiding the use of original usernames to preserve privacy, and propose alternatives to using original text messages. Results show that response selection relies more on the textual content of conversations, while addressee recognition requires capturing their structural dimension. Using an LLM in a zero-shot setting, we further highlight how sensitivity to prompt variations is task-dependent.
title Do LLMs suffer from Multi-Party Hangover? A Diagnostic Approach to Addressee Recognition and Response Selection in Conversations
topic Computation and Language
url https://arxiv.org/abs/2409.18602