Do We Still Need Humans in the Loop? Comparing Human and LLM Annotation in Active Learning for Hostility Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hakimi, Ahmad Dawar, Hirlimann, Lea, Augenstein, Isabelle, Schütze, Hinrich |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
From Human Annotation to Automation: LLM-in-the-Loop Active Learning for Arabic Sentiment Analysis
von: Refai, Dania, et al.
Veröffentlicht: (2025)
von: Refai, Dania, et al.
Veröffentlicht: (2025)
MemeScouts@LT-EDI 2026: Asking the Right Questions -- Prompted Weak Supervision for Meme Hate Speech Detection
von: Bueno, Ivo, et al.
Veröffentlicht: (2026)
von: Bueno, Ivo, et al.
Veröffentlicht: (2026)
Time Course MechInterp: Analyzing the Evolution of Components and Knowledge in Large Language Models
von: Hakimi, Ahmad Dawar, et al.
Veröffentlicht: (2025)
von: Hakimi, Ahmad Dawar, et al.
Veröffentlicht: (2025)
On Relation-Specific Neurons in Large Language Models
von: Liu, Yihong, et al.
Veröffentlicht: (2025)
von: Liu, Yihong, et al.
Veröffentlicht: (2025)
SLAyiNG: Towards Queer Language Processing
von: Veloso, Leonor, et al.
Veröffentlicht: (2025)
von: Veloso, Leonor, et al.
Veröffentlicht: (2025)
Relational Linearity is a Predictor of Hallucinations
von: Lu, Yuetian, et al.
Veröffentlicht: (2026)
von: Lu, Yuetian, et al.
Veröffentlicht: (2026)
Exploring the Role of Transliteration in In-Context Learning for Low-resource Languages Written in Non-Latin Scripts
von: Ma, Chunlan, et al.
Veröffentlicht: (2024)
von: Ma, Chunlan, et al.
Veröffentlicht: (2024)
A Federated Approach to Few-Shot Hate Speech Detection for Marginalized Communities
von: Ye, Haotian, et al.
Veröffentlicht: (2024)
von: Ye, Haotian, et al.
Veröffentlicht: (2024)
GlotCC: An Open Broad-Coverage CommonCrawl Corpus and Pipeline for Minority Languages
von: Kargaran, Amir Hossein, et al.
Veröffentlicht: (2024)
von: Kargaran, Amir Hossein, et al.
Veröffentlicht: (2024)
Comparing LLM Text Annotation Skills: A Study on Human Rights Violations in Social Media Data
von: Nemkova, Poli Apollinaire, et al.
Veröffentlicht: (2025)
von: Nemkova, Poli Apollinaire, et al.
Veröffentlicht: (2025)
Measuring and Benchmarking Large Language Models' Capabilities to Generate Persuasive Language
von: Pauli, Amalie Brogaard, et al.
Veröffentlicht: (2024)
von: Pauli, Amalie Brogaard, et al.
Veröffentlicht: (2024)
MoSECroT: Model Stitching with Static Word Embeddings for Crosslingual Zero-shot Transfer
von: Ye, Haotian, et al.
Veröffentlicht: (2024)
von: Ye, Haotian, et al.
Veröffentlicht: (2024)
Enhancing Text Classification through LLM-Driven Active Learning and Human Annotation
von: Rouzegar, Hamidreza, et al.
Veröffentlicht: (2024)
von: Rouzegar, Hamidreza, et al.
Veröffentlicht: (2024)
Analysing Differences in Persuasive Language in LLM-Generated Text: Uncovering Stereotypical Gender Patterns
von: Pauli, Amalie Brogaard, et al.
Veröffentlicht: (2026)
von: Pauli, Amalie Brogaard, et al.
Veröffentlicht: (2026)
Presumed Cultural Identity: How Names Shape LLM Responses
von: Pawar, Siddhesh, et al.
Veröffentlicht: (2025)
von: Pawar, Siddhesh, et al.
Veröffentlicht: (2025)
Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment
von: Atasoy, I. F., et al.
Veröffentlicht: (2026)
von: Atasoy, I. F., et al.
Veröffentlicht: (2026)
Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need?
von: Kim, Jane Paik
Veröffentlicht: (2026)
von: Kim, Jane Paik
Veröffentlicht: (2026)
Can Community Notes Replace Professional Fact-Checkers?
von: Borenstein, Nadav, et al.
Veröffentlicht: (2025)
von: Borenstein, Nadav, et al.
Veröffentlicht: (2025)
ImpliRet: Benchmarking the Implicit Fact Retrieval Challenge
von: Taghavi, Zeinab Sadat, et al.
Veröffentlicht: (2025)
von: Taghavi, Zeinab Sadat, et al.
Veröffentlicht: (2025)
BMIKE-53: Investigating Cross-Lingual Knowledge Editing with In-Context Learning
von: Nie, Ercong, et al.
Veröffentlicht: (2024)
von: Nie, Ercong, et al.
Veröffentlicht: (2024)
Can We Still Hear the Accent? Investigating the Resilience of Native Language Signals in the LLM Era
von: Utami, Nabelanita, et al.
Veröffentlicht: (2026)
von: Utami, Nabelanita, et al.
Veröffentlicht: (2026)
Show Me the Work: Fact-Checkers' Requirements for Explainable Automated Fact-Checking
von: Warren, Greta, et al.
Veröffentlicht: (2025)
von: Warren, Greta, et al.
Veröffentlicht: (2025)
Stress Testing Factual Consistency Metrics for Long-Document Summarization
von: Mujahid, Zain Muhammad, et al.
Veröffentlicht: (2025)
von: Mujahid, Zain Muhammad, et al.
Veröffentlicht: (2025)
Multi-Step Knowledge Interaction Analysis via Rank-2 Subspace Disentanglement
von: Islam, Sekh Mainul, et al.
Veröffentlicht: (2025)
von: Islam, Sekh Mainul, et al.
Veröffentlicht: (2025)
Do Theory of Mind Benchmarks Need Explicit Human-like Reasoning in Language Models?
von: Lu, Yi-Long, et al.
Veröffentlicht: (2025)
von: Lu, Yi-Long, et al.
Veröffentlicht: (2025)
GNNavi: Navigating the Information Flow in Large Language Models by Graph Neural Network
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2024)
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2024)
Enhancing Robustness of Autoregressive Language Models against Orthographic Attacks via Pixel-based Approach
von: Yang, Han, et al.
Veröffentlicht: (2025)
von: Yang, Han, et al.
Veröffentlicht: (2025)
Hybrid Human-LLM Corpus Construction and LLM Evaluation for Rare Linguistic Phenomena
von: Weissweiler, Leonie, et al.
Veröffentlicht: (2024)
von: Weissweiler, Leonie, et al.
Veröffentlicht: (2024)
LongForm: Effective Instruction Tuning with Reverse Instructions
von: Köksal, Abdullatif, et al.
Veröffentlicht: (2023)
von: Köksal, Abdullatif, et al.
Veröffentlicht: (2023)
CRAFT Your Dataset: Task-Specific Synthetic Dataset Generation Through Corpus Retrieval and Augmentation
von: Ziegler, Ingo, et al.
Veröffentlicht: (2024)
von: Ziegler, Ingo, et al.
Veröffentlicht: (2024)
Semantic Sensitivities and Inconsistent Predictions: Measuring the Fragility of NLI Models
von: Arakelyan, Erik, et al.
Veröffentlicht: (2024)
von: Arakelyan, Erik, et al.
Veröffentlicht: (2024)
Text Annotation via Inductive Coding: Comparing Human Experts to LLMs in Qualitative Data Analysis
von: Parfenova, Angelina, et al.
Veröffentlicht: (2025)
von: Parfenova, Angelina, et al.
Veröffentlicht: (2025)
Claim Verification in the Age of Large Language Models: A Survey
von: Dmonte, Alphaeus, et al.
Veröffentlicht: (2024)
von: Dmonte, Alphaeus, et al.
Veröffentlicht: (2024)
Do We Need Frontier Models to Verify Mathematical Proofs?
von: Naik, Aaditya, et al.
Veröffentlicht: (2026)
von: Naik, Aaditya, et al.
Veröffentlicht: (2026)
Why Do We Laugh? Annotation and Taxonomy Generation for Laughable Contexts in Spontaneous Text Conversation
von: Inoue, Koji, et al.
Veröffentlicht: (2025)
von: Inoue, Koji, et al.
Veröffentlicht: (2025)
SynDARin: Synthesising Datasets for Automated Reasoning in Low-Resource Languages
von: Ghazaryan, Gayane, et al.
Veröffentlicht: (2024)
von: Ghazaryan, Gayane, et al.
Veröffentlicht: (2024)
How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2025)
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2025)
Revealing the Parametric Knowledge of Language Models: A Unified Framework for Attribution Methods
von: Yu, Haeun, et al.
Veröffentlicht: (2024)
von: Yu, Haeun, et al.
Veröffentlicht: (2024)
Large Language Models Are Effective Human Annotation Assistants, But Not Good Independent Annotators
von: Gu, Feng, et al.
Veröffentlicht: (2025)
von: Gu, Feng, et al.
Veröffentlicht: (2025)
Human and LLM Biases in Hate Speech Annotations: A Socio-Demographic Analysis of Annotators and Targets
von: Giorgi, Tommaso, et al.
Veröffentlicht: (2024)
von: Giorgi, Tommaso, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
From Human Annotation to Automation: LLM-in-the-Loop Active Learning for Arabic Sentiment Analysis
von: Refai, Dania, et al.
Veröffentlicht: (2025) -
MemeScouts@LT-EDI 2026: Asking the Right Questions -- Prompted Weak Supervision for Meme Hate Speech Detection
von: Bueno, Ivo, et al.
Veröffentlicht: (2026) -
Time Course MechInterp: Analyzing the Evolution of Components and Knowledge in Large Language Models
von: Hakimi, Ahmad Dawar, et al.
Veröffentlicht: (2025) -
On Relation-Specific Neurons in Large Language Models
von: Liu, Yihong, et al.
Veröffentlicht: (2025) -
SLAyiNG: Towards Queer Language Processing
von: Veloso, Leonor, et al.
Veröffentlicht: (2025)