Overview of Dialog System Evaluation Track: Dimensionality, Language, Culture and Safety at DSTC 12

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mendonça, John, Zhang, Lining, Mallidi, Rahul, Lavie, Alon, Trancoso, Isabel, D'Haro, Luis Fernando, Sedoc, João
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915497912565760
author Mendonça, John
Zhang, Lining
Mallidi, Rahul
Lavie, Alon
Trancoso, Isabel
D'Haro, Luis Fernando
Sedoc, João
author_facet Mendonça, John
Zhang, Lining
Mallidi, Rahul
Lavie, Alon
Trancoso, Isabel
D'Haro, Luis Fernando
Sedoc, João
contents The rapid advancement of Large Language Models (LLMs) has intensified the need for robust dialogue system evaluation, yet comprehensive assessment remains challenging. Traditional metrics often prove insufficient, and safety considerations are frequently narrowly defined or culturally biased. The DSTC12 Track 1, "Dialog System Evaluation: Dimensionality, Language, Culture and Safety," is part of the ongoing effort to address these critical gaps. The track comprised two subtasks: (1) Dialogue-level, Multi-dimensional Automatic Evaluation Metrics, and (2) Multilingual and Multicultural Safety Detection. For Task 1, focused on 10 dialogue dimensions, a Llama-3-8B baseline achieved the highest average Spearman's correlation (0.1681), indicating substantial room for improvement. In Task 2, while participating teams significantly outperformed a Llama-Guard-3-1B baseline on the multilingual safety subset (top ROC-AUC 0.9648), the baseline proved superior on the cultural subset (0.5126 ROC-AUC), highlighting critical needs in culturally-aware safety. This paper describes the datasets and baselines provided to participants, as well as submission evaluation results for each of the two proposed subtasks.
format Preprint
id arxiv_https___arxiv_org_abs_2509_13569
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Overview of Dialog System Evaluation Track: Dimensionality, Language, Culture and Safety at DSTC 12
Mendonça, John
Zhang, Lining
Mallidi, Rahul
Lavie, Alon
Trancoso, Isabel
D'Haro, Luis Fernando
Sedoc, João
Computation and Language
The rapid advancement of Large Language Models (LLMs) has intensified the need for robust dialogue system evaluation, yet comprehensive assessment remains challenging. Traditional metrics often prove insufficient, and safety considerations are frequently narrowly defined or culturally biased. The DSTC12 Track 1, "Dialog System Evaluation: Dimensionality, Language, Culture and Safety," is part of the ongoing effort to address these critical gaps. The track comprised two subtasks: (1) Dialogue-level, Multi-dimensional Automatic Evaluation Metrics, and (2) Multilingual and Multicultural Safety Detection. For Task 1, focused on 10 dialogue dimensions, a Llama-3-8B baseline achieved the highest average Spearman's correlation (0.1681), indicating substantial room for improvement. In Task 2, while participating teams significantly outperformed a Llama-Guard-3-1B baseline on the multilingual safety subset (top ROC-AUC 0.9648), the baseline proved superior on the cultural subset (0.5126 ROC-AUC), highlighting critical needs in culturally-aware safety. This paper describes the datasets and baselines provided to participants, as well as submission evaluation results for each of the two proposed subtasks.
title Overview of Dialog System Evaluation Track: Dimensionality, Language, Culture and Safety at DSTC 12
topic Computation and Language
url https://arxiv.org/abs/2509.13569