Human-Annotated NER Dataset for the Kyrgyz Language
Fuente:
arXiv
Saved in:
| Main Authors: | Turatali, Timur, Alekseev, Anton, Jumalieva, Gulira, Kabaeva, Gulnara, Nikolenko, Sergey |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HJ-Ky-0.1: an Evaluation Dataset for Kyrgyz Word Embeddings
by: Alekseev, Anton, et al.
Published: (2024)
by: Alekseev, Anton, et al.
Published: (2024)
KyrgyzNLP: Challenges, Progress, and Future
by: Alekseev, Anton, et al.
Published: (2024)
by: Alekseev, Anton, et al.
Published: (2024)
Syntactic Transfer to Kyrgyz Using the Treebank Translation Method
by: Alekseev, Anton, et al.
Published: (2024)
by: Alekseev, Anton, et al.
Published: (2024)
KyrgyzBERT: A Compact, Efficient Language Model for Kyrgyz NLP
by: Metinov, Adilet, et al.
Published: (2025)
by: Metinov, Adilet, et al.
Published: (2025)
Theoretical Foundations of GPU-Native Compilation for Rapid Code Iteration
by: Metinov, Adilet, et al.
Published: (2025)
by: Metinov, Adilet, et al.
Published: (2025)
Augmenting NER Datasets with LLMs: Towards Automated and Refined Annotation
by: Naraki, Yuji, et al.
Published: (2024)
by: Naraki, Yuji, et al.
Published: (2024)
Adaptive Soft Rolling KV Freeze with Entropy-Guided Recovery: Sublinear Memory Growth for Efficient LLM Inference
by: Metinov, Adilet, et al.
Published: (2025)
by: Metinov, Adilet, et al.
Published: (2025)
The GELATO Dataset for Legislative NER
by: Flynn, Matthew, et al.
Published: (2026)
by: Flynn, Matthew, et al.
Published: (2026)
Toolken+: Improving LLM Tool Usage with Reranking and a Reject Option
by: Yakovlev, Konstantin, et al.
Published: (2024)
by: Yakovlev, Konstantin, et al.
Published: (2024)
Searching by Code: a New SearchBySnippet Dataset and SnippeR Retrieval Model for Searching by Code Snippets
by: Sedykh, Ivan, et al.
Published: (2023)
by: Sedykh, Ivan, et al.
Published: (2023)
Annotation Errors and NER: A Study with OntoNotes 5.0
by: Bernier-Colborne, Gabriel, et al.
Published: (2024)
by: Bernier-Colborne, Gabriel, et al.
Published: (2024)
Label Unification for Cross-Dataset Generalization in Cybersecurity NER
by: Jalocha, Maciej, et al.
Published: (2025)
by: Jalocha, Maciej, et al.
Published: (2025)
CCT-Code: Cross-Consistency Training for Multilingual Clone Detection and Code Search
by: Tikhonov, Anton, et al.
Published: (2023)
by: Tikhonov, Anton, et al.
Published: (2023)
VerifiNER: Verification-augmented NER via Knowledge-grounded Reasoning with Large Language Models
by: Kim, Seoyeon, et al.
Published: (2024)
by: Kim, Seoyeon, et al.
Published: (2024)
On-the-fly Definition Augmentation of LLMs for Biomedical NER
by: Munnangi, Monica, et al.
Published: (2024)
by: Munnangi, Monica, et al.
Published: (2024)
Towards DS-NER: Unveiling and Addressing Latent Noise in Distant Annotations
by: Ding, Yuyang, et al.
Published: (2025)
by: Ding, Yuyang, et al.
Published: (2025)
NERsocial: Efficient Named Entity Recognition Dataset Construction for Human-Robot Interaction Utilizing RapidNER
by: Atuhurra, Jesse, et al.
Published: (2024)
by: Atuhurra, Jesse, et al.
Published: (2024)
OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ Languages
by: Palen-Michel, Chester, et al.
Published: (2024)
by: Palen-Michel, Chester, et al.
Published: (2024)
CMNER: A Chinese Multimodal NER Dataset based on Social Media
by: Ji, Yuanze, et al.
Published: (2024)
by: Ji, Yuanze, et al.
Published: (2024)
2M-NER: Contrastive Learning for Multilingual and Multimodal NER with Language and Modal Fusion
by: Wang, Dongsheng, et al.
Published: (2024)
by: Wang, Dongsheng, et al.
Published: (2024)
L3Cube-MahaSocialNER: A Social Media based Marathi NER Dataset and BERT models
by: Chaudhari, Harsh, et al.
Published: (2023)
by: Chaudhari, Harsh, et al.
Published: (2023)
BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages
by: Muhammad, Shamsuddeen Hassan, et al.
Published: (2025)
by: Muhammad, Shamsuddeen Hassan, et al.
Published: (2025)
MariNER: A Dataset for Historical Brazilian Portuguese Named Entity Recognition
by: Sarcinelli, João Lucas Luz Lima, et al.
Published: (2025)
by: Sarcinelli, João Lucas Luz Lima, et al.
Published: (2025)
OpenMed NER: Open-Source, Domain-Adapted State-of-the-Art Transformers for Biomedical NER Across 12 Public Datasets
by: Panahi, Maziyar
Published: (2025)
by: Panahi, Maziyar
Published: (2025)
Query2Diagram: Answering Developer Queries with UML Diagrams
by: Baryshnikov, Oleg, et al.
Published: (2026)
by: Baryshnikov, Oleg, et al.
Published: (2026)
SLIMER-IT: Zero-Shot NER on Italian Language
by: Zamai, Andrew, et al.
Published: (2024)
by: Zamai, Andrew, et al.
Published: (2024)
YoNER: A New Yorùbá Multi-domain Named Entity Recognition Dataset
by: Falola, Peace Busola, et al.
Published: (2026)
by: Falola, Peace Busola, et al.
Published: (2026)
PrionNER: A Named Entity Recognition Dataset for Prion Disease Biomedical Literature
by: Dao, An, et al.
Published: (2026)
by: Dao, An, et al.
Published: (2026)
Biomedical Nested NER with Large Language Model and UMLS Heuristics
by: Zhou, Wenxin
Published: (2024)
by: Zhou, Wenxin
Published: (2024)
HQP: A Human-Annotated Dataset for Detecting Online Propaganda
by: Maarouf, Abdurahman, et al.
Published: (2023)
by: Maarouf, Abdurahman, et al.
Published: (2023)
Retrieval Augmented Instruction Tuning for Open NER with Large Language Models
by: Xie, Tingyu, et al.
Published: (2024)
by: Xie, Tingyu, et al.
Published: (2024)
ANCHOLIK-NER: A Benchmark Dataset for Bangla Regional Named Entity Recognition
by: Paul, Bidyarthi, et al.
Published: (2025)
by: Paul, Bidyarthi, et al.
Published: (2025)
FiNER-ORD: Financial Named Entity Recognition Open Research Dataset
by: Shah, Agam, et al.
Published: (2023)
by: Shah, Agam, et al.
Published: (2023)
NuNER: Entity Recognition Encoder Pre-training via LLM-Annotated Data
by: Bogdanov, Sergei, et al.
Published: (2024)
by: Bogdanov, Sergei, et al.
Published: (2024)
The Million-Label NER: Breaking Scale Barriers with GLiNER bi-encoder
by: Stepanov, Ihor, et al.
Published: (2026)
by: Stepanov, Ihor, et al.
Published: (2026)
WikiNER-fr-gold: A Gold-Standard NER Corpus
by: Cao, Danrun, et al.
Published: (2024)
by: Cao, Danrun, et al.
Published: (2024)
Astro-NER -- Astronomy Named Entity Recognition: Is GPT a Good Domain Expert Annotator?
by: Evans, Julia, et al.
Published: (2024)
by: Evans, Julia, et al.
Published: (2024)
Tokenization Matters: Improving Zero-Shot NER for Indic Languages
by: Pattnayak, Priyaranjan, et al.
Published: (2025)
by: Pattnayak, Priyaranjan, et al.
Published: (2025)
PET: An Annotated Dataset for Process Extraction from Natural Language Text
by: Bellan, Patrizio, et al.
Published: (2022)
by: Bellan, Patrizio, et al.
Published: (2022)
Do LLMs Surpass Encoders for Biomedical NER?
by: Obeidat, Motasem S, et al.
Published: (2025)
by: Obeidat, Motasem S, et al.
Published: (2025)
Similar Items
-
HJ-Ky-0.1: an Evaluation Dataset for Kyrgyz Word Embeddings
by: Alekseev, Anton, et al.
Published: (2024) -
KyrgyzNLP: Challenges, Progress, and Future
by: Alekseev, Anton, et al.
Published: (2024) -
Syntactic Transfer to Kyrgyz Using the Treebank Translation Method
by: Alekseev, Anton, et al.
Published: (2024) -
KyrgyzBERT: A Compact, Efficient Language Model for Kyrgyz NLP
by: Metinov, Adilet, et al.
Published: (2025) -
Theoretical Foundations of GPU-Native Compilation for Rapid Code Iteration
by: Metinov, Adilet, et al.
Published: (2025)