SPOT: An Annotated French Corpus and Benchmark for Detecting Critical Interventions in Online Conversations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Berriche, Manon, Nouri, Célia, Clavel, Chloée, Cointet, Jean-Philippe
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912953409732608
author Berriche, Manon
Nouri, Célia
Clavel, Chloée
Cointet, Jean-Philippe
author_facet Berriche, Manon
Nouri, Célia
Clavel, Chloée
Cointet, Jean-Philippe
contents We introduce SPOT (Stopping Points in Online Threads), the first annotated corpus translating the sociological concept of stopping point into a reproducible NLP task. Stopping points are ordinary critical interventions that pause or redirect online discussions through a range of forms (irony, subtle doubt or fragmentary arguments) that frameworks like counterspeech or social correction often overlook. We operationalize this concept as a binary classification task and provide reliable annotation guidelines. The corpus contains 43,305 manually annotated French Facebook comments linked to URLs flagged as false information by social media users, enriched with contextual metadata (article, post, parent comment, page or group, and source). We benchmark fine-tuned encoder models (CamemBERT) and instruction-tuned LLMs under various prompting strategies. Results show that fine-tuned encoders outperform prompted LLMs in F1 score by more than 10 percentage points, confirming the importance of supervised learning for emerging non-English social media tasks. Incorporating contextual metadata further improves encoder models F1 scores from 0.75 to 0.78. We release the anonymized dataset, along with the annotation guidelines and code in our code repository, to foster transparency and reproducible research.
format Preprint
id arxiv_https___arxiv_org_abs_2511_07405
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SPOT: An Annotated French Corpus and Benchmark for Detecting Critical Interventions in Online Conversations
Berriche, Manon
Nouri, Célia
Clavel, Chloée
Cointet, Jean-Philippe
Computation and Language
Computers and Society
We introduce SPOT (Stopping Points in Online Threads), the first annotated corpus translating the sociological concept of stopping point into a reproducible NLP task. Stopping points are ordinary critical interventions that pause or redirect online discussions through a range of forms (irony, subtle doubt or fragmentary arguments) that frameworks like counterspeech or social correction often overlook. We operationalize this concept as a binary classification task and provide reliable annotation guidelines. The corpus contains 43,305 manually annotated French Facebook comments linked to URLs flagged as false information by social media users, enriched with contextual metadata (article, post, parent comment, page or group, and source). We benchmark fine-tuned encoder models (CamemBERT) and instruction-tuned LLMs under various prompting strategies. Results show that fine-tuned encoders outperform prompted LLMs in F1 score by more than 10 percentage points, confirming the importance of supervised learning for emerging non-English social media tasks. Incorporating contextual metadata further improves encoder models F1 scores from 0.75 to 0.78. We release the anonymized dataset, along with the annotation guidelines and code in our code repository, to foster transparency and reproducible research.
title SPOT: An Annotated French Corpus and Benchmark for Detecting Critical Interventions in Online Conversations
topic Computation and Language
Computers and Society
url https://arxiv.org/abs/2511.07405