Automatically Labeling Clinical Trial Outcomes: A Large-Scale Benchmark for Drug Development
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Chufan, Pradeepkumar, Jathurshan, Das, Trisha, Thati, Shivashankar, Sun, Jimeng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Neural Signals Generate Clinical Notes in the Wild
by: Pradeepkumar, Jathurshan, et al.
Published: (2026)
by: Pradeepkumar, Jathurshan, et al.
Published: (2026)
$\texttt{AMEND++}$: Benchmarking Eligibility Criteria Amendments in Clinical Trials
by: Das, Trisha, et al.
Published: (2026)
by: Das, Trisha, et al.
Published: (2026)
Test-Time Adaptation for EEG Foundation Models: A Systematic Study under Real-World Distribution Shifts
by: Lee, Gabriel Jason, et al.
Published: (2026)
by: Lee, Gabriel Jason, et al.
Published: (2026)
Synthetic Patient-Physician Dialogue Generation from Clinical Notes Using LLM
by: Das, Trisha, et al.
Published: (2024)
by: Das, Trisha, et al.
Published: (2024)
Tokenizing Single-Channel EEG with Time-Frequency Motif Learning
by: Pradeepkumar, Jathurshan, et al.
Published: (2025)
by: Pradeepkumar, Jathurshan, et al.
Published: (2025)
Prostate-VarBench: A Benchmark with Interpretable TabNet Framework for Prostate Cancer Variant Classification
by: Tavara, Abraham Francisco Arellano, et al.
Published: (2025)
by: Tavara, Abraham Francisco Arellano, et al.
Published: (2025)
SECRET: Semi-supervised Clinical Trial Document Similarity Search
by: Das, Trisha, et al.
Published: (2025)
by: Das, Trisha, et al.
Published: (2025)
Language Interaction Network for Clinical Trial Approval Estimation
by: Gao, Chufan, et al.
Published: (2024)
by: Gao, Chufan, et al.
Published: (2024)
Developing Large Language Models for Clinical Research Using One Million Clinical Trials
by: Wang, Zifeng, et al.
Published: (2025)
by: Wang, Zifeng, et al.
Published: (2025)
Making Conformal Predictors Robust in Healthcare Settings: a Case Study on EEG Classification
by: Chatterjee, Arjun, et al.
Published: (2026)
by: Chatterjee, Arjun, et al.
Published: (2026)
MediTab: Scaling Medical Tabular Data Predictors via Data Consolidation, Enrichment, and Refinement
by: Wang, Zifeng, et al.
Published: (2023)
by: Wang, Zifeng, et al.
Published: (2023)
Towards Adapting Open-Source Large Language Models for Expert-Level Clinical Note Generation
by: Wang, Hanyin, et al.
Published: (2024)
by: Wang, Hanyin, et al.
Published: (2024)
TTM-RE: Memory-Augmented Document-Level Relation Extraction
by: Gao, Chufan, et al.
Published: (2024)
by: Gao, Chufan, et al.
Published: (2024)
MindMap: Knowledge Graph Prompting Sparks Graph of Thoughts in Large Language Models
by: Wen, Yilin, et al.
Published: (2023)
by: Wen, Yilin, et al.
Published: (2023)
PyHealth 2.0: A Comprehensive Open-Source Toolkit for Accessible and Reproducible Clinical Deep Learning
by: Wu, John, et al.
Published: (2026)
by: Wu, John, et al.
Published: (2026)
TrialSynth: Generation of Synthetic Sequential Clinical Trial Data
by: Gao, Chufan, et al.
Published: (2024)
by: Gao, Chufan, et al.
Published: (2024)
Comprehensive Reassessment of Large-Scale Evaluation Outcomes in LLMs: A Multifaceted Statistical Approach
by: Sun, Kun, et al.
Published: (2024)
by: Sun, Kun, et al.
Published: (2024)
Automatic Labelling with Open-source LLMs using Dynamic Label Schema Integration
by: Walshe, Thomas, et al.
Published: (2025)
by: Walshe, Thomas, et al.
Published: (2025)
CTBench: A Comprehensive Benchmark for Evaluating Language Model Capabilities in Clinical Trial Design
by: Neehal, Nafis, et al.
Published: (2024)
by: Neehal, Nafis, et al.
Published: (2024)
ERBench: An Entity-Relationship based Automatically Verifiable Hallucination Benchmark for Large Language Models
by: Oh, Jio, et al.
Published: (2024)
by: Oh, Jio, et al.
Published: (2024)
RDMA: Cost Effective Agent-Driven Rare Disease Mining from Electronic Health Records
by: Wu, John, et al.
Published: (2025)
by: Wu, John, et al.
Published: (2025)
A Large-Scale Benchmark for Evaluating Large Language Models on Medical Question Answering in Romanian
by: Rogoz, Ana-Cristina, et al.
Published: (2025)
by: Rogoz, Ana-Cristina, et al.
Published: (2025)
Enhancing Hepatopathy Clinical Trial Efficiency: A Secure, Large Language Model-Powered Pre-Screening Pipeline
by: Gui, Xiongbin, et al.
Published: (2025)
by: Gui, Xiongbin, et al.
Published: (2025)
EpiGraph: Building Generalists for Evidence-Intensive Epilepsy Reasoning in the Wild
by: Dai, Yuyang, et al.
Published: (2026)
by: Dai, Yuyang, et al.
Published: (2026)
MLB: A Scenario-Driven Benchmark for Evaluating Large Language Models in Clinical Applications
by: He, Qing, et al.
Published: (2026)
by: He, Qing, et al.
Published: (2026)
An Experimental Design Framework for Label-Efficient Supervised Finetuning of Large Language Models
by: Bhatt, Gantavya, et al.
Published: (2024)
by: Bhatt, Gantavya, et al.
Published: (2024)
Revisiting the Role of Label Smoothing in Enhanced Text Sentiment Classification
by: Gao, Yijie, et al.
Published: (2023)
by: Gao, Yijie, et al.
Published: (2023)
Large Language Models in Drug Discovery and Development: From Disease Mechanisms to Clinical Trials
by: Zheng, Yizhen, et al.
Published: (2024)
by: Zheng, Yizhen, et al.
Published: (2024)
How Well Do Multimodal Models Reason on ECG Signals?
by: Xu, Maxwell A., et al.
Published: (2026)
by: Xu, Maxwell A., et al.
Published: (2026)
Transformer-based Joint Modelling for Automatic Essay Scoring and Off-Topic Detection
by: Das, Sourya Dipta, et al.
Published: (2024)
by: Das, Sourya Dipta, et al.
Published: (2024)
Scaling BERT Models for Turkish Automatic Punctuation and Capitalization Correction
by: Saoud, Abdulkader, et al.
Published: (2024)
by: Saoud, Abdulkader, et al.
Published: (2024)
MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes
by: Abacha, Asma Ben, et al.
Published: (2024)
by: Abacha, Asma Ben, et al.
Published: (2024)
ALHD: A Large-Scale and Multigenre Benchmark Dataset for Arabic LLM-Generated Text Detection
by: Khairallah, Ali, et al.
Published: (2025)
by: Khairallah, Ali, et al.
Published: (2025)
ReEval: Automatic Hallucination Evaluation for Retrieval-Augmented Large Language Models via Transferable Adversarial Attacks
by: Yu, Xiaodong, et al.
Published: (2023)
by: Yu, Xiaodong, et al.
Published: (2023)
ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation
by: Liu, Qin, et al.
Published: (2025)
by: Liu, Qin, et al.
Published: (2025)
Deep Learning-based Prediction of Clinical Trial Enrollment with Uncertainty Estimates
by: Do, Tien Huu, et al.
Published: (2025)
by: Do, Tien Huu, et al.
Published: (2025)
Evaluating the Factuality of Large Language Models using Large-Scale Knowledge Graphs
by: Liu, Xiaoze, et al.
Published: (2024)
by: Liu, Xiaoze, et al.
Published: (2024)
Benchmarking Benchmark Leakage in Large Language Models
by: Xu, Ruijie, et al.
Published: (2024)
by: Xu, Ruijie, et al.
Published: (2024)
NoiseBench: Benchmarking the Impact of Real Label Noise on Named Entity Recognition
by: Merdjanovska, Elena, et al.
Published: (2024)
by: Merdjanovska, Elena, et al.
Published: (2024)
PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models
by: Xu, Yinggan, et al.
Published: (2025)
by: Xu, Yinggan, et al.
Published: (2025)
Similar Items
-
Neural Signals Generate Clinical Notes in the Wild
by: Pradeepkumar, Jathurshan, et al.
Published: (2026) -
$\texttt{AMEND++}$: Benchmarking Eligibility Criteria Amendments in Clinical Trials
by: Das, Trisha, et al.
Published: (2026) -
Test-Time Adaptation for EEG Foundation Models: A Systematic Study under Real-World Distribution Shifts
by: Lee, Gabriel Jason, et al.
Published: (2026) -
Synthetic Patient-Physician Dialogue Generation from Clinical Notes Using LLM
by: Das, Trisha, et al.
Published: (2024) -
Tokenizing Single-Channel EEG with Time-Frequency Motif Learning
by: Pradeepkumar, Jathurshan, et al.
Published: (2025)