FC-CONAN: An Exhaustively Paired Dataset for Robust Evaluation of Retrieval Systems
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866915707262861312 |
|---|---|
| author | Junqueras, Juan Boudin, Florian Zin, May-Myo Nguyen, Ha-Thanh Fungwacharakorn, Wachara Furman, Damián Ariel Aizawa, Akiko Satoh, Ken |
| author_facet | Junqueras, Juan Boudin, Florian Zin, May-Myo Nguyen, Ha-Thanh Fungwacharakorn, Wachara Furman, Damián Ariel Aizawa, Akiko Satoh, Ken |
| contents | Hate speech (HS) is a critical issue in online discourse, and one promising strategy to counter it is through the use of counter-narratives (CNs). Datasets linking HS with CNs are essential for advancing counterspeech research. However, even flagship resources like CONAN (Chung et al., 2019) annotate only a sparse subset of all possible HS-CN pairs, limiting evaluation. We introduce FC-CONAN (Fully Connected CONAN), the first dataset created by exhaustively considering all combinations of 45 English HS messages and 129 CNs. A two-stage annotation process involving nine annotators and four validators produces four partitions-Diamond, Gold, Silver, and Bronze-that balance reliability and scale. None of the labeled pairs overlap with CONAN, uncovering hundreds of previously unlabelled positives. FC-CONAN enables more faithful evaluation of counterspeech retrieval systems and facilitates detailed error analysis. The dataset is publicly available. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_01350 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | FC-CONAN: An Exhaustively Paired Dataset for Robust Evaluation of Retrieval Systems Junqueras, Juan Boudin, Florian Zin, May-Myo Nguyen, Ha-Thanh Fungwacharakorn, Wachara Furman, Damián Ariel Aizawa, Akiko Satoh, Ken Computation and Language Hate speech (HS) is a critical issue in online discourse, and one promising strategy to counter it is through the use of counter-narratives (CNs). Datasets linking HS with CNs are essential for advancing counterspeech research. However, even flagship resources like CONAN (Chung et al., 2019) annotate only a sparse subset of all possible HS-CN pairs, limiting evaluation. We introduce FC-CONAN (Fully Connected CONAN), the first dataset created by exhaustively considering all combinations of 45 English HS messages and 129 CNs. A two-stage annotation process involving nine annotators and four validators produces four partitions-Diamond, Gold, Silver, and Bronze-that balance reliability and scale. None of the labeled pairs overlap with CONAN, uncovering hundreds of previously unlabelled positives. FC-CONAN enables more faithful evaluation of counterspeech retrieval systems and facilitates detailed error analysis. The dataset is publicly available. |
| title | FC-CONAN: An Exhaustively Paired Dataset for Robust Evaluation of Retrieval Systems |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2601.01350 |