Do We Need Language-Specific Fact-Checking Models? The Case of Chinese

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Caiqi, Guo, Zhijiang, Vlachos, Andreas
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909335020371968
author Zhang, Caiqi
Guo, Zhijiang
Vlachos, Andreas
author_facet Zhang, Caiqi
Guo, Zhijiang
Vlachos, Andreas
contents This paper investigates the potential benefits of language-specific fact-checking models, focusing on the case of Chinese. We first demonstrate the limitations of translation-based methods and multilingual large language models (e.g., GPT-4), highlighting the need for language-specific systems. We further propose a Chinese fact-checking system that can better retrieve evidence from a document by incorporating context information. To better analyze token-level biases in different systems, we construct an adversarial dataset based on the CHEF dataset, where each instance has large word overlap with the original one but holds the opposite veracity label. Experimental results on the CHEF dataset and our adversarial dataset show that our proposed method outperforms translation-based methods and multilingual LLMs and is more robust toward biases, while there is still large room for improvement, emphasizing the importance of language-specific fact-checking systems.
format Preprint
id arxiv_https___arxiv_org_abs_2401_15498
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Do We Need Language-Specific Fact-Checking Models? The Case of Chinese
Zhang, Caiqi
Guo, Zhijiang
Vlachos, Andreas
Computation and Language
This paper investigates the potential benefits of language-specific fact-checking models, focusing on the case of Chinese. We first demonstrate the limitations of translation-based methods and multilingual large language models (e.g., GPT-4), highlighting the need for language-specific systems. We further propose a Chinese fact-checking system that can better retrieve evidence from a document by incorporating context information. To better analyze token-level biases in different systems, we construct an adversarial dataset based on the CHEF dataset, where each instance has large word overlap with the original one but holds the opposite veracity label. Experimental results on the CHEF dataset and our adversarial dataset show that our proposed method outperforms translation-based methods and multilingual LLMs and is more robust toward biases, while there is still large room for improvement, emphasizing the importance of language-specific fact-checking systems.
title Do We Need Language-Specific Fact-Checking Models? The Case of Chinese
topic Computation and Language
url https://arxiv.org/abs/2401.15498