BioCoref: Benchmarking Biomedical Coreference Resolution with LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Salem, Nourah M, White, Elizabeth, Bada, Michael, Hunter, Lawrence |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GAPMAP: Mapping Scientific Knowledge Gaps in Biomedical Literature Using Large Language Models
by: Salem, Nourah M, et al.
Published: (2025)
by: Salem, Nourah M, et al.
Published: (2025)
Exploring Multiple Strategies to Improve Multilingual Coreference Resolution in CorefUD
by: Pražák, Ondřej, et al.
Published: (2024)
by: Pražák, Ondřej, et al.
Published: (2024)
CorefInst: Leveraging LLMs for Multilingual Coreference Resolution
by: Arslan, Tuğba Pamay, et al.
Published: (2025)
by: Arslan, Tuğba Pamay, et al.
Published: (2025)
ThaiCoref: Thai Coreference Resolution Dataset
by: Trakuekul, Pontakorn, et al.
Published: (2024)
by: Trakuekul, Pontakorn, et al.
Published: (2024)
ImCoref-CeS: An Improved Lightweight Pipeline for Coreference Resolution with LLM-based Checker-Splitter Refinement
by: Luo, Kangyang, et al.
Published: (2025)
by: Luo, Kangyang, et al.
Published: (2025)
Do Lexical and Contextual Coreference Resolution Systems Degrade Differently under Mention Noise? An Empirical Study on Scientific Software Mentions
by: Alkan, Atilla Kaan, et al.
Published: (2026)
by: Alkan, Atilla Kaan, et al.
Published: (2026)
Major Entity Identification: A Generalizable Alternative to Coreference Resolution
by: Manikantan, Kawshik, et al.
Published: (2024)
by: Manikantan, Kawshik, et al.
Published: (2024)
MediSwift: Efficient Sparse Pre-trained Biomedical Language Models
by: Thangarasa, Vithursan, et al.
Published: (2024)
by: Thangarasa, Vithursan, et al.
Published: (2024)
BioNCERE: Non-Contrastive Enhancement For Relation Extraction In Biomedical Texts
by: Noravesh, Farshad
Published: (2024)
by: Noravesh, Farshad
Published: (2024)
BioPars: A Pretrained Biomedical Large Language Model for Persian Biomedical Text Mining
by: Merzah, Baqer M., et al.
Published: (2025)
by: Merzah, Baqer M., et al.
Published: (2025)
Inside CORE-KG: Evaluating Structured Prompting and Coreference Resolution for Knowledge Graphs
by: Meher, Dipak, et al.
Published: (2025)
by: Meher, Dipak, et al.
Published: (2025)
Improving LLMs' Learning for Coreference Resolution
by: Gan, Yujian, et al.
Published: (2025)
by: Gan, Yujian, et al.
Published: (2025)
Bias in LLMs as Annotators: The Effect of Party Cues on Labelling Decision by Large Language Models
by: Vera, Sebastian Vallejo, et al.
Published: (2024)
by: Vera, Sebastian Vallejo, et al.
Published: (2024)
Who Are All The Stochastic Parrots Imitating? They Should Tell Us!
by: Shaier, Sagi, et al.
Published: (2023)
by: Shaier, Sagi, et al.
Published: (2023)
CaresAI at BioCreative IX Track 1 -- LLM for Biomedical QA
by: Abdel-Salam, Reem, et al.
Published: (2025)
by: Abdel-Salam, Reem, et al.
Published: (2025)
LLMs are not Zero-Shot Reasoners for Biomedical Information Extraction
by: Nagar, Aishik, et al.
Published: (2024)
by: Nagar, Aishik, et al.
Published: (2024)
Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content
by: Touchent, Rian, et al.
Published: (2025)
by: Touchent, Rian, et al.
Published: (2025)
Comparing Template-based and Template-free Language Model Probing
by: Shaier, Sagi, et al.
Published: (2024)
by: Shaier, Sagi, et al.
Published: (2024)
DrBenchmark: A Large Language Understanding Evaluation Benchmark for French Biomedical Domain
by: Labrak, Yanis, et al.
Published: (2024)
by: Labrak, Yanis, et al.
Published: (2024)
Benchmarking LLMs' Judgments with No Gold Standard
by: Xu, Shengwei, et al.
Published: (2024)
by: Xu, Shengwei, et al.
Published: (2024)
Benchmarking and Understanding Compositional Relational Reasoning of LLMs
by: Ni, Ruikang, et al.
Published: (2024)
by: Ni, Ruikang, et al.
Published: (2024)
PATENTWRITER: A Benchmarking Study for Patent Drafting with LLMs
by: Shomee, Homaira Huda, et al.
Published: (2025)
by: Shomee, Homaira Huda, et al.
Published: (2025)
Odysseys: Benchmarking Web Agents on Realistic Long Horizon Tasks
by: Jang, Lawrence Keunho, et al.
Published: (2026)
by: Jang, Lawrence Keunho, et al.
Published: (2026)
Beyond Retrieval: Ensembling Cross-Encoders and GPT Rerankers with LLMs for Biomedical QA
by: Verma, Shashank, et al.
Published: (2025)
by: Verma, Shashank, et al.
Published: (2025)
MUCAR: Benchmarking Multilingual Cross-Modal Ambiguity Resolution for Multimodal Large Language Models
by: Wang, Xiaolong, et al.
Published: (2025)
by: Wang, Xiaolong, et al.
Published: (2025)
Are Smaller Open-Weight LLMs Closing the Gap to Proprietary Models for Biomedical Question Answering?
by: Stachura, Damian, et al.
Published: (2025)
by: Stachura, Damian, et al.
Published: (2025)
Tabular LLMs for Interpretable Few-Shot Alzheimer's Disease Prediction with Multimodal Biomedical Data
by: Kearney, Sophie, et al.
Published: (2026)
by: Kearney, Sophie, et al.
Published: (2026)
PORT: Preference Optimization on Reasoning Traces
by: Lahlou, Salem, et al.
Published: (2024)
by: Lahlou, Salem, et al.
Published: (2024)
CRCE: Coreference-Retention Concept Erasure in Text-to-Image Diffusion Models
by: Xue, Yuyang, et al.
Published: (2025)
by: Xue, Yuyang, et al.
Published: (2025)
Evaluating LLMs' Multilingual Capabilities for Bengali: Benchmark Creation and Performance Analysis
by: Bhowmik, Shimanto, et al.
Published: (2025)
by: Bhowmik, Shimanto, et al.
Published: (2025)
XFinBench: Benchmarking LLMs in Complex Financial Problem Solving and Reasoning
by: Zhang, Zhihan, et al.
Published: (2025)
by: Zhang, Zhihan, et al.
Published: (2025)
LabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs
by: Zhou, Yujun, et al.
Published: (2024)
by: Zhou, Yujun, et al.
Published: (2024)
IdentifyMe: A Challenging Long-Context Mention Resolution Benchmark for LLMs
by: Manikantan, Kawshik, et al.
Published: (2024)
by: Manikantan, Kawshik, et al.
Published: (2024)
Benchmarking Hindi LLMs: A New Suite of Datasets and a Comparative Analysis
by: Kamath, Anusha, et al.
Published: (2025)
by: Kamath, Anusha, et al.
Published: (2025)
WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communications
by: Li, Xin, et al.
Published: (2025)
by: Li, Xin, et al.
Published: (2025)
MPIB: A Benchmark for Medical Prompt Injection Attacks and Clinical Safety in LLMs
by: Lee, Junhyeok, et al.
Published: (2026)
by: Lee, Junhyeok, et al.
Published: (2026)
PeruMedQA: Benchmarking Large Language Models (LLMs) on Peruvian Medical Exams -- Dataset Construction and Evaluation
by: Carrillo-Larco, Rodrigo M., et al.
Published: (2025)
by: Carrillo-Larco, Rodrigo M., et al.
Published: (2025)
GRILE: A Benchmark for Grammar Reasoning and Explanation in Romanian LLMs
by: Dumitran, Adrian-Marius, et al.
Published: (2025)
by: Dumitran, Adrian-Marius, et al.
Published: (2025)
BioCoder: A Benchmark for Bioinformatics Code Generation with Large Language Models
by: Tang, Xiangru, et al.
Published: (2023)
by: Tang, Xiangru, et al.
Published: (2023)
WikiContradict: A Benchmark for Evaluating LLMs on Real-World Knowledge Conflicts from Wikipedia
by: Hou, Yufang, et al.
Published: (2024)
by: Hou, Yufang, et al.
Published: (2024)
Similar Items
-
GAPMAP: Mapping Scientific Knowledge Gaps in Biomedical Literature Using Large Language Models
by: Salem, Nourah M, et al.
Published: (2025) -
Exploring Multiple Strategies to Improve Multilingual Coreference Resolution in CorefUD
by: Pražák, Ondřej, et al.
Published: (2024) -
CorefInst: Leveraging LLMs for Multilingual Coreference Resolution
by: Arslan, Tuğba Pamay, et al.
Published: (2025) -
ThaiCoref: Thai Coreference Resolution Dataset
by: Trakuekul, Pontakorn, et al.
Published: (2024) -
ImCoref-CeS: An Improved Lightweight Pipeline for Coreference Resolution with LLM-based Checker-Splitter Refinement
by: Luo, Kangyang, et al.
Published: (2025)