Beyond Dataset Creation: Critical View of Annotation Variation and Bias Probing of a Dataset for Online Radical Content Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Riabi, Arij, Mouilleron, Virginie, Mahamdi, Menel, Antoun, Wissam, Seddah, Djamé |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cloaked Classifiers: Pseudonymization Strategies on Sensitive Classification Tasks
by: Riabi, Arij, et al.
Published: (2024)
by: Riabi, Arij, et al.
Published: (2024)
Enriching the NArabizi Treebank: A Multifaceted Approach to Supporting an Under-Resourced Language
by: Riabi, Arij, et al.
Published: (2023)
by: Riabi, Arij, et al.
Published: (2023)
From Text to Source: Results in Detecting Large Language Model-Generated Content
by: Antoun, Wissam, et al.
Published: (2023)
by: Antoun, Wissam, et al.
Published: (2023)
Can Character-based Language Models Improve Downstream Task Performance in Low-Resource and Noisy Language Scenarios?
by: Riabi, Arij, et al.
Published: (2021)
by: Riabi, Arij, et al.
Published: (2021)
Common Ground, Diverse Roots: The Difficulty of Classifying Common Examples in Spanish Varieties
by: Lopetegui, Javier A., et al.
Published: (2024)
by: Lopetegui, Javier A., et al.
Published: (2024)
ModernBERT or DeBERTaV3? Examining Architecture and Data Influence on Transformer Encoder Models Performance
by: Antoun, Wissam, et al.
Published: (2025)
by: Antoun, Wissam, et al.
Published: (2025)
When Tables Go Crazy: Evaluating Multimodal Models on French Financial Documents
by: Mouilleron, Virginie, et al.
Published: (2026)
by: Mouilleron, Virginie, et al.
Published: (2026)
Triggers Hijack Language Circuits: A Mechanistic Analysis of Backdoor Behaviors in Large Language Models
by: Lasnier, Théo, et al.
Published: (2026)
by: Lasnier, Théo, et al.
Published: (2026)
Multilingual, Multimodal Pipeline for Creating Authentic and Structured Fact-Checked Claim Dataset
by: Hüsünbeyi, Z. Melce, et al.
Published: (2026)
by: Hüsünbeyi, Z. Melce, et al.
Published: (2026)
Language-Switching Triggers Take a Latent Detour Through Language Models
by: Kulumba, Francis, et al.
Published: (2026)
by: Kulumba, Francis, et al.
Published: (2026)
CamemBERT 2.0: A Smarter French Language Model Aged to Perfection
by: Antoun, Wissam, et al.
Published: (2024)
by: Antoun, Wissam, et al.
Published: (2024)
Leveraging Wikidata for Geographically Informed Sociocultural Bias Dataset Creation: Application to Latin America
by: Karmim, Yannis, et al.
Published: (2026)
by: Karmim, Yannis, et al.
Published: (2026)
Gaperon: A Peppered English-French Generative Language Model Suite
by: Godey, Nathan, et al.
Published: (2025)
by: Godey, Nathan, et al.
Published: (2025)
Rethinking the Multilingual Reasoning Gap with Layer Swap
by: Lasbordes, Maxence, et al.
Published: (2026)
by: Lasbordes, Maxence, et al.
Published: (2026)
Disentangling meaning from language in LLM-based machine translation
by: Lasnier, Théo, et al.
Published: (2026)
by: Lasnier, Théo, et al.
Published: (2026)
BiaSWE: An Expert Annotated Dataset for Misogyny Detection in Swedish
by: Kukk, Kätriin, et al.
Published: (2025)
by: Kukk, Kätriin, et al.
Published: (2025)
HQP: A Human-Annotated Dataset for Detecting Online Propaganda
by: Maarouf, Abdurahman, et al.
Published: (2023)
by: Maarouf, Abdurahman, et al.
Published: (2023)
IYKYK: Using language models to decode extremist cryptolects
by: de Kock, Christine, et al.
Published: (2025)
by: de Kock, Christine, et al.
Published: (2025)
The Promises and Pitfalls of LLM Annotations in Dataset Labeling: a Case Study on Media Bias Detection
by: Horych, Tomas, et al.
Published: (2024)
by: Horych, Tomas, et al.
Published: (2024)
RuBia: A Russian Language Bias Detection Dataset
by: Grigoreva, Veronika, et al.
Published: (2024)
by: Grigoreva, Veronika, et al.
Published: (2024)
Dataset Creation and Baseline Models for Sexism Detection in Hausa
by: Muhammad, Fatima Adam, et al.
Published: (2025)
by: Muhammad, Fatima Adam, et al.
Published: (2025)
HALvest-Contrastive: Retrieval-Like Authorship Attribution with Patch-Level Late Interaction
by: Kulumba, Francis, et al.
Published: (2024)
by: Kulumba, Francis, et al.
Published: (2024)
Sina at FigNews 2024: Multilingual Datasets Annotated with Bias and Propaganda
by: Duaibes, Lina, et al.
Published: (2024)
by: Duaibes, Lina, et al.
Published: (2024)
The MediaSpin Dataset: Post-Publication News Headline Edits Annotated for Media Bias
by: Verma, Preetika, et al.
Published: (2024)
by: Verma, Preetika, et al.
Published: (2024)
Cost-aware LLM-based Online Dataset Annotation
by: Elumar, Eray Can, et al.
Published: (2025)
by: Elumar, Eray Can, et al.
Published: (2025)
An Annotated Dataset of Errors in Premodern Greek and Baselines for Detecting Them
by: Brooks, Creston, et al.
Published: (2024)
by: Brooks, Creston, et al.
Published: (2024)
Dataset Creation for Visual Entailment using Generative AI
by: Reijtenbach, Rob, et al.
Published: (2025)
by: Reijtenbach, Rob, et al.
Published: (2025)
SPOT: An Annotated French Corpus and Benchmark for Detecting Critical Interventions in Online Conversations
by: Berriche, Manon, et al.
Published: (2025)
by: Berriche, Manon, et al.
Published: (2025)
L-ReLF: A Framework for Lexical Dataset Creation
by: Sedrati, Anass, et al.
Published: (2026)
by: Sedrati, Anass, et al.
Published: (2026)
A Highly Clean Recipe Dataset with Ingredient States Annotation for State Probing Task
by: Toyooka, Mashiro, et al.
Published: (2025)
by: Toyooka, Mashiro, et al.
Published: (2025)
Human-Annotated NER Dataset for the Kyrgyz Language
by: Turatali, Timur, et al.
Published: (2025)
by: Turatali, Timur, et al.
Published: (2025)
Annotation Tool and Dataset for Fact-Checking Podcasts
by: Setty, Vinay, et al.
Published: (2025)
by: Setty, Vinay, et al.
Published: (2025)
Analyzing Dataset Annotation Quality Management in the Wild
by: Klie, Jan-Christoph, et al.
Published: (2023)
by: Klie, Jan-Christoph, et al.
Published: (2023)
Donkii: Can Annotation Error Detection Methods Find Errors in Instruction-Tuning Datasets?
by: Weber-Genzel, Leon, et al.
Published: (2023)
by: Weber-Genzel, Leon, et al.
Published: (2023)
Creation of the Estonian Subjectivity Dataset: Assessing the Degree of Subjectivity on a Scale
by: Gailit, Karl Gustav, et al.
Published: (2025)
by: Gailit, Karl Gustav, et al.
Published: (2025)
Modeling Romanized Hindi and Bengali: Dataset Creation and Multilingual LLM Integration
by: Gharami, Kanchon, et al.
Published: (2025)
by: Gharami, Kanchon, et al.
Published: (2025)
SACRED: A Faithful Annotated Multimedia Multimodal Multilingual Dataset for Classifying Connectedness Types in Online Spirituality
by: Guan, Qinghao, et al.
Published: (2026)
by: Guan, Qinghao, et al.
Published: (2026)
Beyond the Explicit: A Bilingual Dataset for Dehumanization Detection in Social Media
by: Assenmacher, Dennis, et al.
Published: (2025)
by: Assenmacher, Dennis, et al.
Published: (2025)
Enhancing Adverse Drug Event Detection with Multimodal Dataset: Corpus Creation and Model Development
by: Sahoo, Pranab, et al.
Published: (2024)
by: Sahoo, Pranab, et al.
Published: (2024)
Synthetic Dataset Creation and Fine-Tuning of Transformer Models for Question Answering in Serbian
by: Cvetanović, Aleksa, et al.
Published: (2024)
by: Cvetanović, Aleksa, et al.
Published: (2024)
Similar Items
-
Cloaked Classifiers: Pseudonymization Strategies on Sensitive Classification Tasks
by: Riabi, Arij, et al.
Published: (2024) -
Enriching the NArabizi Treebank: A Multifaceted Approach to Supporting an Under-Resourced Language
by: Riabi, Arij, et al.
Published: (2023) -
From Text to Source: Results in Detecting Large Language Model-Generated Content
by: Antoun, Wissam, et al.
Published: (2023) -
Can Character-based Language Models Improve Downstream Task Performance in Low-Resource and Noisy Language Scenarios?
by: Riabi, Arij, et al.
Published: (2021) -
Common Ground, Diverse Roots: The Difficulty of Classifying Common Examples in Spanish Varieties
by: Lopetegui, Javier A., et al.
Published: (2024)