CaseReportBench: An LLM Benchmark Dataset for Dense Information Extraction in Clinical Case Reports
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Xiao Yu Cindy, Ferreira, Carlos R., Rossignol, Francis, Ng, Raymond T., Wasserman, Wyeth, Zhu, Jian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ELMTEX: Fine-Tuning Large Language Models for Structured Clinical Information Extraction. A Case Study on Clinical Reports
von: Guluzade, Aynur, et al.
Veröffentlicht: (2025)
von: Guluzade, Aynur, et al.
Veröffentlicht: (2025)
Filling in the Clinical Gaps in Benchmark: Case for HealthBench for the Japanese medical system
von: Hisada, Shohei, et al.
Veröffentlicht: (2025)
von: Hisada, Shohei, et al.
Veröffentlicht: (2025)
Converting Annotated Clinical Cases into Structured Case Report Forms
von: Ferrazzi, Pietro, et al.
Veröffentlicht: (2025)
von: Ferrazzi, Pietro, et al.
Veröffentlicht: (2025)
Low-resource Information Extraction with the European Clinical Case Corpus
von: Ghosh, Soumitra, et al.
Veröffentlicht: (2025)
von: Ghosh, Soumitra, et al.
Veröffentlicht: (2025)
BURExtract-Llama: An LLM for Clinical Concept Extraction in Breast Ultrasound Reports
von: Chen, Yuxuan, et al.
Veröffentlicht: (2024)
von: Chen, Yuxuan, et al.
Veröffentlicht: (2024)
AppealCase: A Dataset and Benchmark for Civil Case Appeal Scenarios
von: Huang, Yuting, et al.
Veröffentlicht: (2025)
von: Huang, Yuting, et al.
Veröffentlicht: (2025)
LLM-based Triplet Extraction from Financial Reports
von: Wesslund, Dante, et al.
Veröffentlicht: (2026)
von: Wesslund, Dante, et al.
Veröffentlicht: (2026)
Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench
von: Perlitz, Yotam, et al.
Veröffentlicht: (2024)
von: Perlitz, Yotam, et al.
Veröffentlicht: (2024)
ShoppingBench: A Real-World Intent-Grounded Shopping Benchmark for LLM-based Agents
von: Wang, Jiangyuan, et al.
Veröffentlicht: (2025)
von: Wang, Jiangyuan, et al.
Veröffentlicht: (2025)
LexRel: Benchmarking Legal Relation Extraction for Chinese Civil Cases
von: Cai, Yida, et al.
Veröffentlicht: (2025)
von: Cai, Yida, et al.
Veröffentlicht: (2025)
LLM, Reporting In! Medical Information Extraction Across Prompting, Fine-tuning and Post-correction
von: Belmadani, Ikram, et al.
Veröffentlicht: (2025)
von: Belmadani, Ikram, et al.
Veröffentlicht: (2025)
Inductive Bias Extraction and Matching for LLM Prompts
von: Angel, Christian M., et al.
Veröffentlicht: (2025)
von: Angel, Christian M., et al.
Veröffentlicht: (2025)
Key Coverage Matters: Semi-Structured Extraction of OCR Clinical Reports
von: Wang, Yu, et al.
Veröffentlicht: (2026)
von: Wang, Yu, et al.
Veröffentlicht: (2026)
ESG-Bench: Benchmarking Long-Context ESG Reports for Hallucination Mitigation
von: Sun, Siqi, et al.
Veröffentlicht: (2026)
von: Sun, Siqi, et al.
Veröffentlicht: (2026)
Causal Tree Extraction from Medical Case Reports: A Novel Task for Experts-like Text Comprehension
von: Yahata, Sakiko, et al.
Veröffentlicht: (2025)
von: Yahata, Sakiko, et al.
Veröffentlicht: (2025)
A Large-Language Model Framework for Relative Timeline Extraction from PubMed Case Reports
von: Wang, Jing, et al.
Veröffentlicht: (2025)
von: Wang, Jing, et al.
Veröffentlicht: (2025)
EffiBench-X: A Multi-Language Benchmark for Measuring Efficiency of LLM-Generated Code
von: Qing, Yuhao, et al.
Veröffentlicht: (2025)
von: Qing, Yuhao, et al.
Veröffentlicht: (2025)
AIR-Bench: Automated Heterogeneous Information Retrieval Benchmark
von: Chen, Jianlyu, et al.
Veröffentlicht: (2024)
von: Chen, Jianlyu, et al.
Veröffentlicht: (2024)
OI-Bench: An Option Injection Benchmark for Evaluating LLM Susceptibility to Directive Interference
von: Liou, Yow-Fu, et al.
Veröffentlicht: (2026)
von: Liou, Yow-Fu, et al.
Veröffentlicht: (2026)
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
von: Deng, Shihan, et al.
Veröffentlicht: (2024)
von: Deng, Shihan, et al.
Veröffentlicht: (2024)
ER-Reason: A Benchmark Dataset for LLM Clinical Reasoning in the Emergency Room
von: Mehandru, Nikita, et al.
Veröffentlicht: (2025)
von: Mehandru, Nikita, et al.
Veröffentlicht: (2025)
DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions
von: Wang, Xinran, et al.
Veröffentlicht: (2026)
von: Wang, Xinran, et al.
Veröffentlicht: (2026)
$\textit{BenchIE}^{FL}$ : A Manually Re-Annotated Fact-Based Open Information Extraction Benchmark
von: Lamarche, Fabrice, et al.
Veröffentlicht: (2024)
von: Lamarche, Fabrice, et al.
Veröffentlicht: (2024)
LLM-Powered Grapheme-to-Phoneme Conversion: Benchmark and Case Study
von: Qharabagh, Mahta Fetrat, et al.
Veröffentlicht: (2024)
von: Qharabagh, Mahta Fetrat, et al.
Veröffentlicht: (2024)
Who Benchmarks the Benchmarks? A Case Study of LLM Evaluation in Icelandic
von: Ingimundarson, Finnur Ágúst, et al.
Veröffentlicht: (2026)
von: Ingimundarson, Finnur Ágúst, et al.
Veröffentlicht: (2026)
FlowBench: Revisiting and Benchmarking Workflow-Guided Planning for LLM-based Agents
von: Xiao, Ruixuan, et al.
Veröffentlicht: (2024)
von: Xiao, Ruixuan, et al.
Veröffentlicht: (2024)
CUPCase: Clinically Uncommon Patient Cases and Diagnoses Dataset
von: Perets, Oriel, et al.
Veröffentlicht: (2025)
von: Perets, Oriel, et al.
Veröffentlicht: (2025)
AlpsBench: An LLM Personalization Benchmark for Real-Dialogue Memorization and Preference Alignment
von: Xiao, Jianfei, et al.
Veröffentlicht: (2026)
von: Xiao, Jianfei, et al.
Veröffentlicht: (2026)
LCTG Bench: LLM Controlled Text Generation Benchmark
von: Kurihara, Kentaro, et al.
Veröffentlicht: (2025)
von: Kurihara, Kentaro, et al.
Veröffentlicht: (2025)
MedRepBench: A Comprehensive Benchmark for Medical Report Interpretation
von: Shang, Fangxin, et al.
Veröffentlicht: (2025)
von: Shang, Fangxin, et al.
Veröffentlicht: (2025)
BenchBench: Benchmarking Automated Benchmark Generation
von: Zheng, Yandan, et al.
Veröffentlicht: (2026)
von: Zheng, Yandan, et al.
Veröffentlicht: (2026)
VocalBench-DF: A Benchmark for Evaluating Speech LLM Robustness to Disfluency
von: Liu, Hongcheng, et al.
Veröffentlicht: (2025)
von: Liu, Hongcheng, et al.
Veröffentlicht: (2025)
WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the Wild
von: Zhang, Linhao, et al.
Veröffentlicht: (2025)
von: Zhang, Linhao, et al.
Veröffentlicht: (2025)
GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration
von: Zhao, Junjie, et al.
Veröffentlicht: (2026)
von: Zhao, Junjie, et al.
Veröffentlicht: (2026)
ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences
von: Nguyen, Bang, et al.
Veröffentlicht: (2026)
von: Nguyen, Bang, et al.
Veröffentlicht: (2026)
ClinicalBench: Can LLMs Beat Traditional ML Models in Clinical Prediction?
von: Chen, Canyu, et al.
Veröffentlicht: (2024)
von: Chen, Canyu, et al.
Veröffentlicht: (2024)
EmoBench-UA: A Benchmark Dataset for Emotion Detection in Ukrainian
von: Dementieva, Daryna, et al.
Veröffentlicht: (2025)
von: Dementieva, Daryna, et al.
Veröffentlicht: (2025)
MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
von: Li, Yingyun, et al.
Veröffentlicht: (2026)
von: Li, Yingyun, et al.
Veröffentlicht: (2026)
Efficiently Identifying Low-Quality Language Subsets in Multilingual Datasets: A Case Study on a Large-Scale Multilingual Audio Dataset
von: Samir, Farhan, et al.
Veröffentlicht: (2024)
von: Samir, Farhan, et al.
Veröffentlicht: (2024)
QualBench: Benchmarking Chinese LLMs with Localized Professional Qualifications for Vertical Domain Evaluation
von: Hong, Mengze, et al.
Veröffentlicht: (2025)
von: Hong, Mengze, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ELMTEX: Fine-Tuning Large Language Models for Structured Clinical Information Extraction. A Case Study on Clinical Reports
von: Guluzade, Aynur, et al.
Veröffentlicht: (2025) -
Filling in the Clinical Gaps in Benchmark: Case for HealthBench for the Japanese medical system
von: Hisada, Shohei, et al.
Veröffentlicht: (2025) -
Converting Annotated Clinical Cases into Structured Case Report Forms
von: Ferrazzi, Pietro, et al.
Veröffentlicht: (2025) -
Low-resource Information Extraction with the European Clinical Case Corpus
von: Ghosh, Soumitra, et al.
Veröffentlicht: (2025) -
BURExtract-Llama: An LLM for Clinical Concept Extraction in Breast Ultrasound Reports
von: Chen, Yuxuan, et al.
Veröffentlicht: (2024)