FC-CONAN: An Exhaustively Paired Dataset for Robust Evaluation of Retrieval Systems

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Junqueras, Juan, Boudin, Florian, Zin, May-Myo, Nguyen, Ha-Thanh, Fungwacharakorn, Wachara, Furman, Damián Ariel, Aizawa, Akiko, Satoh, Ken
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915707262861312
author Junqueras, Juan
Boudin, Florian
Zin, May-Myo
Nguyen, Ha-Thanh
Fungwacharakorn, Wachara
Furman, Damián Ariel
Aizawa, Akiko
Satoh, Ken
author_facet Junqueras, Juan
Boudin, Florian
Zin, May-Myo
Nguyen, Ha-Thanh
Fungwacharakorn, Wachara
Furman, Damián Ariel
Aizawa, Akiko
Satoh, Ken
contents Hate speech (HS) is a critical issue in online discourse, and one promising strategy to counter it is through the use of counter-narratives (CNs). Datasets linking HS with CNs are essential for advancing counterspeech research. However, even flagship resources like CONAN (Chung et al., 2019) annotate only a sparse subset of all possible HS-CN pairs, limiting evaluation. We introduce FC-CONAN (Fully Connected CONAN), the first dataset created by exhaustively considering all combinations of 45 English HS messages and 129 CNs. A two-stage annotation process involving nine annotators and four validators produces four partitions-Diamond, Gold, Silver, and Bronze-that balance reliability and scale. None of the labeled pairs overlap with CONAN, uncovering hundreds of previously unlabelled positives. FC-CONAN enables more faithful evaluation of counterspeech retrieval systems and facilitates detailed error analysis. The dataset is publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2601_01350
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FC-CONAN: An Exhaustively Paired Dataset for Robust Evaluation of Retrieval Systems
Junqueras, Juan
Boudin, Florian
Zin, May-Myo
Nguyen, Ha-Thanh
Fungwacharakorn, Wachara
Furman, Damián Ariel
Aizawa, Akiko
Satoh, Ken
Computation and Language
Hate speech (HS) is a critical issue in online discourse, and one promising strategy to counter it is through the use of counter-narratives (CNs). Datasets linking HS with CNs are essential for advancing counterspeech research. However, even flagship resources like CONAN (Chung et al., 2019) annotate only a sparse subset of all possible HS-CN pairs, limiting evaluation. We introduce FC-CONAN (Fully Connected CONAN), the first dataset created by exhaustively considering all combinations of 45 English HS messages and 129 CNs. A two-stage annotation process involving nine annotators and four validators produces four partitions-Diamond, Gold, Silver, and Bronze-that balance reliability and scale. None of the labeled pairs overlap with CONAN, uncovering hundreds of previously unlabelled positives. FC-CONAN enables more faithful evaluation of counterspeech retrieval systems and facilitates detailed error analysis. The dataset is publicly available.
title FC-CONAN: An Exhaustively Paired Dataset for Robust Evaluation of Retrieval Systems
topic Computation and Language
url https://arxiv.org/abs/2601.01350