Do LLMs Judge Distantly Supervised Named Entity Labels Well? Constructing the JudgeWEL Dataset
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Plum, Alistair, Bernardy, Laura, Ranasinghe, Tharindu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Guided Distant Supervision for Multilingual Relation Extraction Data: Adapting to a New Language
von: Plum, Alistair, et al.
Veröffentlicht: (2024)
von: Plum, Alistair, et al.
Veröffentlicht: (2024)
Text Generation Models for Luxembourgish with Limited Data: A Balanced Multilingual Strategy
von: Plum, Alistair, et al.
Veröffentlicht: (2024)
von: Plum, Alistair, et al.
Veröffentlicht: (2024)
SANTA: Separate Strategies for Inaccurate and Incomplete Annotation Noise in Distantly-Supervised Named Entity Recognition
von: Si, Shuzheng, et al.
Veröffentlicht: (2023)
von: Si, Shuzheng, et al.
Veröffentlicht: (2023)
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
von: Thakur, Aman Singh, et al.
Veröffentlicht: (2024)
von: Thakur, Aman Singh, et al.
Veröffentlicht: (2024)
Can LLMs be Good Graph Judge for Knowledge Graph Construction?
von: Huang, Haoyu, et al.
Veröffentlicht: (2024)
von: Huang, Haoyu, et al.
Veröffentlicht: (2024)
Improving the Robustness of Distantly-Supervised Named Entity Recognition via Uncertainty-Aware Teacher Learning and Student-Student Collaborative Learning
von: Si, Shuzheng, et al.
Veröffentlicht: (2023)
von: Si, Shuzheng, et al.
Veröffentlicht: (2023)
SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification
von: Yoon, Kanghoon, et al.
Veröffentlicht: (2025)
von: Yoon, Kanghoon, et al.
Veröffentlicht: (2025)
ALEXSIS-PT: A New Resource for Portuguese Lexical Simplification
von: North, Kai, et al.
Veröffentlicht: (2022)
von: North, Kai, et al.
Veröffentlicht: (2022)
Towards Simulating Social Media Users with LLMs: Evaluating the Operational Validity of Conditioned Comment Prediction
von: Schwager, Nils, et al.
Veröffentlicht: (2026)
von: Schwager, Nils, et al.
Veröffentlicht: (2026)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
von: Shi, Lin, et al.
Veröffentlicht: (2024)
von: Shi, Lin, et al.
Veröffentlicht: (2024)
Towards Generalized Offensive Language Identification
von: Dmonte, Alphaeus, et al.
Veröffentlicht: (2024)
von: Dmonte, Alphaeus, et al.
Veröffentlicht: (2024)
MultiLS: A Multi-task Lexical Simplification Framework
von: North, Kai, et al.
Veröffentlicht: (2024)
von: North, Kai, et al.
Veröffentlicht: (2024)
JudgeLRM: Large Reasoning Models as a Judge
von: Chen, Nuo, et al.
Veröffentlicht: (2025)
von: Chen, Nuo, et al.
Veröffentlicht: (2025)
JudgeLM: Fine-tuned Large Language Models are Scalable Judges
von: Zhu, Lianghui, et al.
Veröffentlicht: (2023)
von: Zhu, Lianghui, et al.
Veröffentlicht: (2023)
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
von: Wang, Yidong, et al.
Veröffentlicht: (2025)
von: Wang, Yidong, et al.
Veröffentlicht: (2025)
LuxInstruct: A Cross-Lingual Instruction Tuning Dataset For Luxembourgish
von: Philippy, Fred, et al.
Veröffentlicht: (2025)
von: Philippy, Fred, et al.
Veröffentlicht: (2025)
Language-Independent Sentiment Labelling with Distant Supervision: A Case Study for English, Sepedi and Setswana
von: Mabokela, Koena Ronny, et al.
Veröffentlicht: (2025)
von: Mabokela, Koena Ronny, et al.
Veröffentlicht: (2025)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
Agent-as-a-Judge
von: You, Runyang, et al.
Veröffentlicht: (2026)
von: You, Runyang, et al.
Veröffentlicht: (2026)
Improving Pseudo Labels with Global-Local Denoising Framework for Cross-lingual Named Entity Recognition
von: Ding, Zhuojun, et al.
Veröffentlicht: (2024)
von: Ding, Zhuojun, et al.
Veröffentlicht: (2024)
EasyJudge: an Easy-to-use Tool for Comprehensive Response Evaluation of LLMs
von: Li, Yijie, et al.
Veröffentlicht: (2024)
von: Li, Yijie, et al.
Veröffentlicht: (2024)
Named Entity Recognition in Historical Italian: The Case of Giacomo Leopardi's Zibaldone
von: Santini, Cristian, et al.
Veröffentlicht: (2025)
von: Santini, Cristian, et al.
Veröffentlicht: (2025)
Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge
von: Sun, Xin, et al.
Veröffentlicht: (2026)
von: Sun, Xin, et al.
Veröffentlicht: (2026)
Exploring the Performance of Large Language Models on Subjective Span Identification Tasks
von: Dmonte, Alphaeus, et al.
Veröffentlicht: (2026)
von: Dmonte, Alphaeus, et al.
Veröffentlicht: (2026)
Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement
von: Han, Steve, et al.
Veröffentlicht: (2025)
von: Han, Steve, et al.
Veröffentlicht: (2025)
ELLEN: Extremely Lightly Supervised Learning For Efficient Named Entity Recognition
von: Riaz, Haris, et al.
Veröffentlicht: (2024)
von: Riaz, Haris, et al.
Veröffentlicht: (2024)
Named Clinical Entity Recognition Benchmark
von: Abdul, Wadood M, et al.
Veröffentlicht: (2024)
von: Abdul, Wadood M, et al.
Veröffentlicht: (2024)
NSINA: A News Corpus for Sinhala
von: Hettiarachchi, Hansi, et al.
Veröffentlicht: (2024)
von: Hettiarachchi, Hansi, et al.
Veröffentlicht: (2024)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
von: Tong, Terry, et al.
Veröffentlicht: (2025)
von: Tong, Terry, et al.
Veröffentlicht: (2025)
LLMs Judge Themselves: A Game-Theoretic Framework for Human-Aligned Evaluation
von: Yang, Gao, et al.
Veröffentlicht: (2025)
von: Yang, Gao, et al.
Veröffentlicht: (2025)
NoiseBench: Benchmarking the Impact of Real Label Noise on Named Entity Recognition
von: Merdjanovska, Elena, et al.
Veröffentlicht: (2024)
von: Merdjanovska, Elena, et al.
Veröffentlicht: (2024)
Named Entity Recognition in COVID-19 tweets with Entity Knowledge Augmentation
von: Zhang, Xuankang, et al.
Veröffentlicht: (2025)
von: Zhang, Xuankang, et al.
Veröffentlicht: (2025)
Summarization Metrics for Spanish and Basque: Do Automatic Scores and LLM-Judges Correlate with Humans?
von: Barnes, Jeremy, et al.
Veröffentlicht: (2025)
von: Barnes, Jeremy, et al.
Veröffentlicht: (2025)
Do Before You Judge: Self-Reference as a Pathway to Better LLM Evaluation
von: Lin, Wei-Hsiang, et al.
Veröffentlicht: (2025)
von: Lin, Wei-Hsiang, et al.
Veröffentlicht: (2025)
JudgeRLVR: Judge First, Generate Second for Efficient Reasoning
von: Duo, Jiangshan, et al.
Veröffentlicht: (2026)
von: Duo, Jiangshan, et al.
Veröffentlicht: (2026)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
Generative Annotation for ASR Named Entity Correction
von: Luo, Yuanchang, et al.
Veröffentlicht: (2025)
von: Luo, Yuanchang, et al.
Veröffentlicht: (2025)
DynClean: Training Dynamics-based Label Cleaning for Distantly-Supervised Named Entity Recognition
von: Zhang, Qi, et al.
Veröffentlicht: (2025)
von: Zhang, Qi, et al.
Veröffentlicht: (2025)
BANER: Boundary-Aware LLMs for Few-Shot Named Entity Recognition
von: Guo, Quanjiang, et al.
Veröffentlicht: (2024)
von: Guo, Quanjiang, et al.
Veröffentlicht: (2024)
Is LLM an Overconfident Judge? Unveiling the Capabilities of LLMs in Detecting Offensive Language with Annotation Disagreement
von: Lu, Junyu, et al.
Veröffentlicht: (2025)
von: Lu, Junyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Guided Distant Supervision for Multilingual Relation Extraction Data: Adapting to a New Language
von: Plum, Alistair, et al.
Veröffentlicht: (2024) -
Text Generation Models for Luxembourgish with Limited Data: A Balanced Multilingual Strategy
von: Plum, Alistair, et al.
Veröffentlicht: (2024) -
SANTA: Separate Strategies for Inaccurate and Incomplete Annotation Noise in Distantly-Supervised Named Entity Recognition
von: Si, Shuzheng, et al.
Veröffentlicht: (2023) -
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
von: Thakur, Aman Singh, et al.
Veröffentlicht: (2024) -
Can LLMs be Good Graph Judge for Knowledge Graph Construction?
von: Huang, Haoyu, et al.
Veröffentlicht: (2024)