Is 'Hope' a person or an idea? A pilot benchmark for NER: comparing traditional NLP tools and large language models on ambiguous entities
Fuente:
arXiv
Saved in:
| Main Author: | Latifi, Payam |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A thorough benchmark of automatic text classification: From traditional approaches to large language models
by: Cunha, Washington, et al.
Published: (2025)
by: Cunha, Washington, et al.
Published: (2025)
Leveraging large language models for efficient representation learning for entity resolution
by: Xu, Xiaowei, et al.
Published: (2024)
by: Xu, Xiaowei, et al.
Published: (2024)
Noise reduction in BERT NER models for clinical entity extraction
by: Jiwani, Kuldeep, et al.
Published: (2026)
by: Jiwani, Kuldeep, et al.
Published: (2026)
Creativity Benchmark: A benchmark for marketing creativity for large language models
by: Bhat, Ninad, et al.
Published: (2025)
by: Bhat, Ninad, et al.
Published: (2025)
A dataset and benchmark for hospital course summarization with adapted large language models
by: Aali, Asad, et al.
Published: (2024)
by: Aali, Asad, et al.
Published: (2024)
Retrieval augmented generation based dynamic prompting for few-shot biomedical named entity recognition using large language models
by: Ge, Yao, et al.
Published: (2025)
by: Ge, Yao, et al.
Published: (2025)
ChildEval: When large language models meet children's personalities
by: Luo, Yanyan, et al.
Published: (2026)
by: Luo, Yanyan, et al.
Published: (2026)
Referential ambiguity and clarification requests: comparing human and LLM behaviour
by: Madge, Chris, et al.
Published: (2025)
by: Madge, Chris, et al.
Published: (2025)
NER- RoBERTa: Fine-Tuning RoBERTa for Named Entity Recognition (NER) within low-resource languages
by: Abdullah, Abdulhady Abas, et al.
Published: (2024)
by: Abdullah, Abdulhady Abas, et al.
Published: (2024)
Batayan: A Filipino NLP benchmark for evaluating Large Language Models
by: Montalan, Jann Railey, et al.
Published: (2025)
by: Montalan, Jann Railey, et al.
Published: (2025)
Dissociating language and thought in large language models
by: Mahowald, Kyle, et al.
Published: (2023)
by: Mahowald, Kyle, et al.
Published: (2023)
LongTail-Swap: benchmarking language models' abilities on rare words
by: Algayres, Robin, et al.
Published: (2025)
by: Algayres, Robin, et al.
Published: (2025)
Context Matters: Comparison of commercial large language tools in veterinary medicine
by: Poore, Tyler J, et al.
Published: (2025)
by: Poore, Tyler J, et al.
Published: (2025)
TelcoLM: collecting data, adapting, and benchmarking language models for the telecommunication domain
by: Barboule, Camille, et al.
Published: (2024)
by: Barboule, Camille, et al.
Published: (2024)
On the attribution of confidence to large language models
by: Keeling, Geoff, et al.
Published: (2024)
by: Keeling, Geoff, et al.
Published: (2024)
WikiNER-fr-gold: A Gold-Standard NER Corpus
by: Cao, Danrun, et al.
Published: (2024)
by: Cao, Danrun, et al.
Published: (2024)
Correcting misinformation on social media with a large language model
by: Zhou, Xinyi, et al.
Published: (2024)
by: Zhou, Xinyi, et al.
Published: (2024)
Evaluating large language models in medical applications: a survey
by: Chen, Xiaolan, et al.
Published: (2024)
by: Chen, Xiaolan, et al.
Published: (2024)
2M-NER: Contrastive Learning for Multilingual and Multimodal NER with Language and Modal Fusion
by: Wang, Dongsheng, et al.
Published: (2024)
by: Wang, Dongsheng, et al.
Published: (2024)
A review on the use of large language models as virtual tutors
by: García-Méndez, Silvia, et al.
Published: (2024)
by: García-Méndez, Silvia, et al.
Published: (2024)
A survey of textual cyber abuse detection using cutting-edge language models and large language models
by: Diaz-Garcia, Jose A., et al.
Published: (2025)
by: Diaz-Garcia, Jose A., et al.
Published: (2025)
PBa-LLM: Privacy- and Bias-aware NLP using Named-Entity Recognition (NER)
by: Mancera, Gonzalo, et al.
Published: (2025)
by: Mancera, Gonzalo, et al.
Published: (2025)
Quantifying non deterministic drift in large language models
by: Nicholson, Claire
Published: (2026)
by: Nicholson, Claire
Published: (2026)
Can large language models build causal graphs?
by: Long, Stephanie, et al.
Published: (2023)
by: Long, Stephanie, et al.
Published: (2023)
Multi-round jailbreak attack on large language models
by: Zhou, Yihua, et al.
Published: (2024)
by: Zhou, Yihua, et al.
Published: (2024)
Response: Emergent analogical reasoning in large language models
by: Hodel, Damian, et al.
Published: (2023)
by: Hodel, Damian, et al.
Published: (2023)
The 20 questions game to distinguish large language models
by: Richardeau, Gurvan, et al.
Published: (2024)
by: Richardeau, Gurvan, et al.
Published: (2024)
Representation in large language models
by: Yetman, Cameron
Published: (2025)
by: Yetman, Cameron
Published: (2025)
Evaluating the performance of state-of-the-art esg domain-specific pre-trained large language models in text classification against existing models and traditional machine learning techniques
by: Chung, Tin Yuet, et al.
Published: (2024)
by: Chung, Tin Yuet, et al.
Published: (2024)
Failure of contextual invariance in large language models
by: Kumar, Sagar, et al.
Published: (2026)
by: Kumar, Sagar, et al.
Published: (2026)
A blind spot for large language models: Supradiegetic linguistic information
by: Zimmerman, Julia Witte, et al.
Published: (2023)
by: Zimmerman, Julia Witte, et al.
Published: (2023)
Hyacinth6B: A large language model for Traditional Chinese
by: Song, Chih-Wei, et al.
Published: (2024)
by: Song, Chih-Wei, et al.
Published: (2024)
ESNERA: Empirical and semantic named entity alignment for named entity dataset merging
by: Zhang, Xiaobo, et al.
Published: (2025)
by: Zhang, Xiaobo, et al.
Published: (2025)
Superhuman performance of a large language model on the reasoning tasks of a physician
by: Brodeur, Peter G., et al.
Published: (2024)
by: Brodeur, Peter G., et al.
Published: (2024)
Streamlining evidence based clinical recommendations with large language models
by: Li, Dubai, et al.
Published: (2025)
by: Li, Dubai, et al.
Published: (2025)
Re-evaluating Theory of Mind evaluation in large language models
by: Hu, Jennifer, et al.
Published: (2025)
by: Hu, Jennifer, et al.
Published: (2025)
Strong and weak alignment of large language models with human values
by: Khamassi, Mehdi, et al.
Published: (2024)
by: Khamassi, Mehdi, et al.
Published: (2024)
Disentangling generalization and memorization in large language models using chess
by: Pleiss, Leonard S., et al.
Published: (2026)
by: Pleiss, Leonard S., et al.
Published: (2026)
MathDivide: Improved mathematical reasoning by large language models
by: Srivastava, Saksham Sahai, et al.
Published: (2024)
by: Srivastava, Saksham Sahai, et al.
Published: (2024)
Can AI grade your essays? A comparative analysis of large language models and teacher ratings in multidimensional essay scoring
by: Seßler, Kathrin, et al.
Published: (2024)
by: Seßler, Kathrin, et al.
Published: (2024)
Similar Items
-
A thorough benchmark of automatic text classification: From traditional approaches to large language models
by: Cunha, Washington, et al.
Published: (2025) -
Leveraging large language models for efficient representation learning for entity resolution
by: Xu, Xiaowei, et al.
Published: (2024) -
Noise reduction in BERT NER models for clinical entity extraction
by: Jiwani, Kuldeep, et al.
Published: (2026) -
Creativity Benchmark: A benchmark for marketing creativity for large language models
by: Bhat, Ninad, et al.
Published: (2025) -
A dataset and benchmark for hospital course summarization with adapted large language models
by: Aali, Asad, et al.
Published: (2024)