Annotation Errors and NER: A Study with OntoNotes 5.0
Fuente:
arXiv
Saved in:
| Main Authors: | Bernier-Colborne, Gabriel, Vajjala, Sowmya |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Test Set Quality in Multilingual LLM Evaluation
by: Kranti, Chalamalasetti, et al.
Published: (2025)
by: Kranti, Chalamalasetti, et al.
Published: (2025)
IndicGEC: Powerful Models, or a Measurement Mirage?
by: Vajjala, Sowmya
Published: (2025)
by: Vajjala, Sowmya
Published: (2025)
The Problem with Safety Classification is not just the Models
by: Vajjala, Sowmya
Published: (2025)
by: Vajjala, Sowmya
Published: (2025)
MetricalARGS: A Taxonomy for Studying Metrical Poetry with LLMs
by: Kranti, Chalamalasetti, et al.
Published: (2025)
by: Kranti, Chalamalasetti, et al.
Published: (2025)
Dravidian language family through Universal Dependencies lens
by: Rama, Taraka, et al.
Published: (2024)
by: Rama, Taraka, et al.
Published: (2024)
Text Classification in the LLM Era -- Where do we stand?
by: Vajjala, Sowmya, et al.
Published: (2025)
by: Vajjala, Sowmya, et al.
Published: (2025)
Does Synthetic Data Help Named Entity Recognition for Low-Resource Languages?
by: Kamath, Gaurav, et al.
Published: (2025)
by: Kamath, Gaurav, et al.
Published: (2025)
MATA: Mindful Assessment of the Telugu Abilities of Large Language Models
by: Kranti, Chalamalasetti, et al.
Published: (2025)
by: Kranti, Chalamalasetti, et al.
Published: (2025)
Scope Ambiguities in Large Language Models
by: Kamath, Gaurav, et al.
Published: (2024)
by: Kamath, Gaurav, et al.
Published: (2024)
LLMs in Education: Novel Perspectives, Challenges, and Opportunities
by: Alhafni, Bashar, et al.
Published: (2024)
by: Alhafni, Bashar, et al.
Published: (2024)
Opportunities and Challenges of LLMs in Education: An NLP Perspective
by: Vajjala, Sowmya, et al.
Published: (2025)
by: Vajjala, Sowmya, et al.
Published: (2025)
Human-Annotated NER Dataset for the Kyrgyz Language
by: Turatali, Timur, et al.
Published: (2025)
by: Turatali, Timur, et al.
Published: (2025)
OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ Languages
by: Palen-Michel, Chester, et al.
Published: (2024)
by: Palen-Michel, Chester, et al.
Published: (2024)
Augmenting NER Datasets with LLMs: Towards Automated and Refined Annotation
by: Naraki, Yuji, et al.
Published: (2024)
by: Naraki, Yuji, et al.
Published: (2024)
Towards DS-NER: Unveiling and Addressing Latent Noise in Distant Annotations
by: Ding, Yuyang, et al.
Published: (2025)
by: Ding, Yuyang, et al.
Published: (2025)
WikiNER-fr-gold: A Gold-Standard NER Corpus
by: Cao, Danrun, et al.
Published: (2024)
by: Cao, Danrun, et al.
Published: (2024)
The GELATO Dataset for Legislative NER
by: Flynn, Matthew, et al.
Published: (2026)
by: Flynn, Matthew, et al.
Published: (2026)
VerifiNER: Verification-augmented NER via Knowledge-grounded Reasoning with Large Language Models
by: Kim, Seoyeon, et al.
Published: (2024)
by: Kim, Seoyeon, et al.
Published: (2024)
NuNER: Entity Recognition Encoder Pre-training via LLM-Annotated Data
by: Bogdanov, Sergei, et al.
Published: (2024)
by: Bogdanov, Sergei, et al.
Published: (2024)
LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories
by: Vishnubhotla, Krishnapriya, et al.
Published: (2026)
by: Vishnubhotla, Krishnapriya, et al.
Published: (2026)
The Million-Label NER: Breaking Scale Barriers with GLiNER bi-encoder
by: Stepanov, Ihor, et al.
Published: (2026)
by: Stepanov, Ihor, et al.
Published: (2026)
Astro-NER -- Astronomy Named Entity Recognition: Is GPT a Good Domain Expert Annotator?
by: Evans, Julia, et al.
Published: (2024)
by: Evans, Julia, et al.
Published: (2024)
ERNIE 5.0 Technical Report
by: Wang, Haifeng, et al.
Published: (2026)
by: Wang, Haifeng, et al.
Published: (2026)
2M-NER: Contrastive Learning for Multilingual and Multimodal NER with Language and Modal Fusion
by: Wang, Dongsheng, et al.
Published: (2024)
by: Wang, Dongsheng, et al.
Published: (2024)
Comparative Analysis of Extrinsic Factors for NER in French
by: Yang, Grace, et al.
Published: (2024)
by: Yang, Grace, et al.
Published: (2024)
On-the-fly Definition Augmentation of LLMs for Biomedical NER
by: Munnangi, Monica, et al.
Published: (2024)
by: Munnangi, Monica, et al.
Published: (2024)
Do LLMs Surpass Encoders for Biomedical NER?
by: Obeidat, Motasem S, et al.
Published: (2025)
by: Obeidat, Motasem S, et al.
Published: (2025)
Novel Benchmark for NER in the Wastewater and Stormwater Domain
by: Cardillo, Franco Alberto, et al.
Published: (2025)
by: Cardillo, Franco Alberto, et al.
Published: (2025)
L3Cube-MahaSocialNER: A Social Media based Marathi NER Dataset and BERT models
by: Chaudhari, Harsh, et al.
Published: (2023)
by: Chaudhari, Harsh, et al.
Published: (2023)
PrOnto: Language Model Evaluations for 859 Languages
by: Gessler, Luke
Published: (2023)
by: Gessler, Luke
Published: (2023)
Comparative Study of Zero-Shot Cross-Lingual Transfer for Bodo POS and NER Tagging Using Gemini 2.0 Flash Thinking Experimental Model
by: Narzary, Sanjib, et al.
Published: (2025)
by: Narzary, Sanjib, et al.
Published: (2025)
ErAConD : Error Annotated Conversational Dialog Dataset for Grammatical Error Correction
by: Yuan, Xun, et al.
Published: (2021)
by: Yuan, Xun, et al.
Published: (2021)
Donkii: Can Annotation Error Detection Methods Find Errors in Instruction-Tuning Datasets?
by: Weber-Genzel, Leon, et al.
Published: (2023)
by: Weber-Genzel, Leon, et al.
Published: (2023)
Label Unification for Cross-Dataset Generalization in Cybersecurity NER
by: Jalocha, Maciej, et al.
Published: (2025)
by: Jalocha, Maciej, et al.
Published: (2025)
Semantic Similarity in Radiology Reports via LLMs and NER
by: Pearson, Beth, et al.
Published: (2025)
by: Pearson, Beth, et al.
Published: (2025)
CMNER: A Chinese Multimodal NER Dataset based on Social Media
by: Ji, Yuanze, et al.
Published: (2024)
by: Ji, Yuanze, et al.
Published: (2024)
HiligayNER: A Baseline Named Entity Recognition Model for Hiligaynon
by: Teves, James Ald, et al.
Published: (2025)
by: Teves, James Ald, et al.
Published: (2025)
An Annotated Dataset of Errors in Premodern Greek and Baselines for Detecting Them
by: Brooks, Creston, et al.
Published: (2024)
by: Brooks, Creston, et al.
Published: (2024)
Marking: Visual Grading with Highlighting Errors and Annotating Missing Bits
by: Sonkar, Shashank, et al.
Published: (2024)
by: Sonkar, Shashank, et al.
Published: (2024)
NER- RoBERTa: Fine-Tuning RoBERTa for Named Entity Recognition (NER) within low-resource languages
by: Abdullah, Abdulhady Abas, et al.
Published: (2024)
by: Abdullah, Abdulhady Abas, et al.
Published: (2024)
Similar Items
-
Test Set Quality in Multilingual LLM Evaluation
by: Kranti, Chalamalasetti, et al.
Published: (2025) -
IndicGEC: Powerful Models, or a Measurement Mirage?
by: Vajjala, Sowmya
Published: (2025) -
The Problem with Safety Classification is not just the Models
by: Vajjala, Sowmya
Published: (2025) -
MetricalARGS: A Taxonomy for Studying Metrical Poetry with LLMs
by: Kranti, Chalamalasetti, et al.
Published: (2025) -
Dravidian language family through Universal Dependencies lens
by: Rama, Taraka, et al.
Published: (2024)