Augmenting NER Datasets with LLMs: Towards Automated and Refined Annotation
Fuente:
arXiv
Saved in:
| Main Authors: | Naraki, Yuji, Yamaki, Ryosuke, Ikeda, Yoshikazu, Horie, Takafumi, Yoshida, Kotaro, Shimizu, Ryotaro, Naganuma, Hiroki |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DisTaC: Conditioning Task Vectors via Distillation for Robust Model Merging
by: Yoshida, Kotaro, et al.
Published: (2025)
by: Yoshida, Kotaro, et al.
Published: (2025)
On Fairness of Task Arithmetic: The Role of Task Vectors
by: Naganuma, Hiroki, et al.
Published: (2025)
by: Naganuma, Hiroki, et al.
Published: (2025)
Towards Understanding Variants of Invariant Risk Minimization through the Lens of Calibration
by: Yoshida, Kotaro, et al.
Published: (2024)
by: Yoshida, Kotaro, et al.
Published: (2024)
Towards DS-NER: Unveiling and Addressing Latent Noise in Distant Annotations
by: Ding, Yuyang, et al.
Published: (2025)
by: Ding, Yuyang, et al.
Published: (2025)
An Empirical Study of Pre-trained Model Selection for Out-of-Distribution Generalization and Calibration
by: Naganuma, Hiroki, et al.
Published: (2023)
by: Naganuma, Hiroki, et al.
Published: (2023)
DiZiNER: Disagreement-guided Instruction Refinement via Pilot Annotation Simulation for Zero-shot Named Entity Recognition
by: Kim, Siun, et al.
Published: (2026)
by: Kim, Siun, et al.
Published: (2026)
L3Cube-MahaSocialNER: A Social Media based Marathi NER Dataset and BERT models
by: Chaudhari, Harsh, et al.
Published: (2023)
by: Chaudhari, Harsh, et al.
Published: (2023)
ANCHOLIK-NER: A Benchmark Dataset for Bangla Regional Named Entity Recognition
by: Paul, Bidyarthi, et al.
Published: (2025)
by: Paul, Bidyarthi, et al.
Published: (2025)
SHED: Shapley-Based Automated Dataset Refinement for Instruction Fine-Tuning
by: He, Yexiao, et al.
Published: (2024)
by: He, Yexiao, et al.
Published: (2024)
The Million-Label NER: Breaking Scale Barriers with GLiNER bi-encoder
by: Stepanov, Ihor, et al.
Published: (2026)
by: Stepanov, Ihor, et al.
Published: (2026)
NuNER: Entity Recognition Encoder Pre-training via LLM-Annotated Data
by: Bogdanov, Sergei, et al.
Published: (2024)
by: Bogdanov, Sergei, et al.
Published: (2024)
Towards Automated Kernel Generation in the Era of LLMs
by: Yu, Yang, et al.
Published: (2026)
by: Yu, Yang, et al.
Published: (2026)
DiLM: Distilling Dataset into Language Model for Text-level Dataset Distillation
by: Maekawa, Aru, et al.
Published: (2024)
by: Maekawa, Aru, et al.
Published: (2024)
Drifting Objectives for Refining Discrete Diffusion Language Models
by: Oba, Daisuke, et al.
Published: (2026)
by: Oba, Daisuke, et al.
Published: (2026)
A Dataset for Pharmacovigilance in German, French, and Japanese: Annotating Adverse Drug Reactions across Languages
by: Raithel, Lisa, et al.
Published: (2024)
by: Raithel, Lisa, et al.
Published: (2024)
Graph Neural Network and NER-Based Text Summarization
by: Khan, Imaad Zaffar, et al.
Published: (2024)
by: Khan, Imaad Zaffar, et al.
Published: (2024)
Revisiting Generalization Measures Beyond IID: An Empirical Study under Distributional Shift
by: Nakai, Sora, et al.
Published: (2026)
by: Nakai, Sora, et al.
Published: (2026)
WhisperNER: Unified Open Named Entity and Speech Recognition
by: Ayache, Gil, et al.
Published: (2024)
by: Ayache, Gil, et al.
Published: (2024)
Static Word Embeddings for Sentence Semantic Representation
by: Wada, Takashi, et al.
Published: (2025)
by: Wada, Takashi, et al.
Published: (2025)
Towards Scalable Automated Alignment of LLMs: A Survey
by: Cao, Boxi, et al.
Published: (2024)
by: Cao, Boxi, et al.
Published: (2024)
Adversarial Demonstration Learning for Low-resource NER Using Dual Similarity
by: Yuan, Guowen, et al.
Published: (2025)
by: Yuan, Guowen, et al.
Published: (2025)
Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs
by: Wang, Qibin, et al.
Published: (2025)
by: Wang, Qibin, et al.
Published: (2025)
Automated Type Annotation in Python Using Large Language Models
by: Bharti, Varun, et al.
Published: (2025)
by: Bharti, Varun, et al.
Published: (2025)
Automated Knowledge Concept Annotation and Question Representation Learning for Knowledge Tracing
by: Ozyurt, Yilmazcan, et al.
Published: (2024)
by: Ozyurt, Yilmazcan, et al.
Published: (2024)
VinePPO: Refining Credit Assignment in RL Training of LLMs
by: Kazemnejad, Amirhossein, et al.
Published: (2024)
by: Kazemnejad, Amirhossein, et al.
Published: (2024)
Re-Examine Distantly Supervised NER: A New Benchmark and a Simple Approach
by: Li, Yuepei, et al.
Published: (2024)
by: Li, Yuepei, et al.
Published: (2024)
Human-Annotated NER Dataset for the Kyrgyz Language
by: Turatali, Timur, et al.
Published: (2025)
by: Turatali, Timur, et al.
Published: (2025)
Semantic Refinement with LLMs for Graph Representations
by: Thapaliya, Safal, et al.
Published: (2025)
by: Thapaliya, Safal, et al.
Published: (2025)
TARDiS : Text Augmentation for Refining Diversity and Separability
by: Kim, Kyungmin, et al.
Published: (2025)
by: Kim, Kyungmin, et al.
Published: (2025)
Refined Direct Preference Optimization with Synthetic Data for Behavioral Alignment of LLMs
by: Gallego, Víctor
Published: (2024)
by: Gallego, Víctor
Published: (2024)
mucAI at WojoodNER 2024: Arabic Named Entity Recognition with Nearest Neighbor Search
by: Abdou, Ahmed, et al.
Published: (2024)
by: Abdou, Ahmed, et al.
Published: (2024)
GLiNER-Relex: A Unified Framework for Joint Named Entity Recognition and Relation Extraction
by: Stepanov, Ihor, et al.
Published: (2026)
by: Stepanov, Ihor, et al.
Published: (2026)
Explaining Black-box Model Predictions via Two-level Nested Feature Attributions with Consistency Property
by: Yoshikawa, Yuya, et al.
Published: (2024)
by: Yoshikawa, Yuya, et al.
Published: (2024)
Reducing Large Language Model Bias with Emphasis on 'Restricted Industries': Automated Dataset Augmentation and Prejudice Quantification
by: Mondal, Devam, et al.
Published: (2024)
by: Mondal, Devam, et al.
Published: (2024)
Knowledge Distillation in Automated Annotation: Supervised Text Classification with LLM-Generated Training Labels
by: Pangakis, Nicholas, et al.
Published: (2024)
by: Pangakis, Nicholas, et al.
Published: (2024)
Iterative Augmentation with Summarization Refinement (IASR) Evaluation for Unstructured Survey data Modeling and Analysis
by: Bhattad, Payal, et al.
Published: (2025)
by: Bhattad, Payal, et al.
Published: (2025)
Evolving LLMs' Self-Refinement Capability via Synergistic Training-Inference Optimization
by: Zeng, Yongcheng, et al.
Published: (2025)
by: Zeng, Yongcheng, et al.
Published: (2025)
MorphNAS: Differentiable Architecture Search for Morphologically-Aware Multilingual NER
by: Devadiga, Prathamesh, et al.
Published: (2025)
by: Devadiga, Prathamesh, et al.
Published: (2025)
Multi-Response Preference Optimization with Augmented Ranking Dataset
by: Gwon, Hansle, et al.
Published: (2024)
by: Gwon, Hansle, et al.
Published: (2024)
Cost-aware LLM-based Online Dataset Annotation
by: Elumar, Eray Can, et al.
Published: (2025)
by: Elumar, Eray Can, et al.
Published: (2025)
Similar Items
-
DisTaC: Conditioning Task Vectors via Distillation for Robust Model Merging
by: Yoshida, Kotaro, et al.
Published: (2025) -
On Fairness of Task Arithmetic: The Role of Task Vectors
by: Naganuma, Hiroki, et al.
Published: (2025) -
Towards Understanding Variants of Invariant Risk Minimization through the Lens of Calibration
by: Yoshida, Kotaro, et al.
Published: (2024) -
Towards DS-NER: Unveiling and Addressing Latent Noise in Distant Annotations
by: Ding, Yuyang, et al.
Published: (2025) -
An Empirical Study of Pre-trained Model Selection for Out-of-Distribution Generalization and Calibration
by: Naganuma, Hiroki, et al.
Published: (2023)