$\texttt{AMEND++}$: Benchmarking Eligibility Criteria Amendments in Clinical Trials
Fuente:
arXiv
Saved in:
| Main Authors: | Das, Trisha, Beigi, Mandis, Aptekar, Jacob, Sun, Jimeng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SECRET: Semi-supervised Clinical Trial Document Similarity Search
by: Das, Trisha, et al.
Published: (2025)
by: Das, Trisha, et al.
Published: (2025)
TrialSynth: Generation of Synthetic Sequential Clinical Trial Data
by: Gao, Chufan, et al.
Published: (2024)
by: Gao, Chufan, et al.
Published: (2024)
SynRL: Aligning Synthetic Clinical Trial Data with Human-preferred Clinical Endpoints Using Reinforcement Learning
by: Das, Trisha, et al.
Published: (2024)
by: Das, Trisha, et al.
Published: (2024)
Automatically Labeling Clinical Trial Outcomes: A Large-Scale Benchmark for Drug Development
by: Gao, Chufan, et al.
Published: (2024)
by: Gao, Chufan, et al.
Published: (2024)
Synthetic Patient-Physician Dialogue Generation from Clinical Notes Using LLM
by: Das, Trisha, et al.
Published: (2024)
by: Das, Trisha, et al.
Published: (2024)
MindMap: Knowledge Graph Prompting Sparks Graph of Thoughts in Large Language Models
by: Wen, Yilin, et al.
Published: (2023)
by: Wen, Yilin, et al.
Published: (2023)
$\texttt{SEM-CTRL}$: Semantically Controlled Decoding
by: Albinhassan, Mohammad, et al.
Published: (2025)
by: Albinhassan, Mohammad, et al.
Published: (2025)
$\texttt{SPECS}$: Faster Test-Time Scaling through Speculative Drafts
by: Cemri, Mert, et al.
Published: (2025)
by: Cemri, Mert, et al.
Published: (2025)
CTBench: A Comprehensive Benchmark for Evaluating Language Model Capabilities in Clinical Trial Design
by: Neehal, Nafis, et al.
Published: (2024)
by: Neehal, Nafis, et al.
Published: (2024)
$\texttt{PatentAgent}$: Intelligent Agent for Automated Pharmaceutical Patent Analysis
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
RDMA: Cost Effective Agent-Driven Rare Disease Mining from Electronic Health Records
by: Wu, John, et al.
Published: (2025)
by: Wu, John, et al.
Published: (2025)
Towards Adapting Open-Source Large Language Models for Expert-Level Clinical Note Generation
by: Wang, Hanyin, et al.
Published: (2024)
by: Wang, Hanyin, et al.
Published: (2024)
Language Interaction Network for Clinical Trial Approval Estimation
by: Gao, Chufan, et al.
Published: (2024)
by: Gao, Chufan, et al.
Published: (2024)
MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes
by: Abacha, Asma Ben, et al.
Published: (2024)
by: Abacha, Asma Ben, et al.
Published: (2024)
Deep Learning-based Prediction of Clinical Trial Enrollment with Uncertainty Estimates
by: Do, Tien Huu, et al.
Published: (2025)
by: Do, Tien Huu, et al.
Published: (2025)
ctELM: Decoding and Manipulating Embeddings of Clinical Trials with Embedding Language Models
by: Ondov, Brian, et al.
Published: (2026)
by: Ondov, Brian, et al.
Published: (2026)
ECG-Reasoning-Benchmark: A Benchmark for Evaluating Clinical Reasoning Capabilities in ECG Interpretation
by: Oh, Jungwoo, et al.
Published: (2026)
by: Oh, Jungwoo, et al.
Published: (2026)
Token-based Decision Criteria Are Suboptimal in In-context Learning
by: Cho, Hakaze, et al.
Published: (2024)
by: Cho, Hakaze, et al.
Published: (2024)
Enhancing Hepatopathy Clinical Trial Efficiency: A Secure, Large Language Model-Powered Pre-Screening Pipeline
by: Gui, Xiongbin, et al.
Published: (2025)
by: Gui, Xiongbin, et al.
Published: (2025)
MedRECT: A Medical Reasoning Benchmark for Error Correction in Clinical Texts
by: Iwase, Naoto, et al.
Published: (2025)
by: Iwase, Naoto, et al.
Published: (2025)
APE: Selective Fine-tuning with Acceptance Criteria for Language Model Adaptation
by: Marín, Javier
Published: (2025)
by: Marín, Javier
Published: (2025)
MIMIC-RD: Can LLMs differentially diagnose rare diseases in real-world clinical settings?
by: AlDin, Zilal Eiz, et al.
Published: (2025)
by: AlDin, Zilal Eiz, et al.
Published: (2025)
MLB: A Scenario-Driven Benchmark for Evaluating Large Language Models in Clinical Applications
by: He, Qing, et al.
Published: (2026)
by: He, Qing, et al.
Published: (2026)
IITK at SemEval-2024 Task 2: Exploring the Capabilities of LLMs for Safe Biomedical Natural Language Inference for Clinical Trials
by: Mandal, Shreyasi, et al.
Published: (2024)
by: Mandal, Shreyasi, et al.
Published: (2024)
Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts
by: Haimes, Jacob, et al.
Published: (2024)
by: Haimes, Jacob, et al.
Published: (2024)
MMLU-SR: A Benchmark for Stress-Testing Reasoning Capability of Large Language Models
by: Wang, Wentian, et al.
Published: (2024)
by: Wang, Wentian, et al.
Published: (2024)
Interactive Benchmarks
by: Yue, Baoqing, et al.
Published: (2026)
by: Yue, Baoqing, et al.
Published: (2026)
ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation
by: Liu, Qin, et al.
Published: (2025)
by: Liu, Qin, et al.
Published: (2025)
The Emperor's New Clothes in Benchmarking? A Rigorous Examination of Mitigation Strategies for LLM Benchmark Data Contamination
by: Sun, Yifan, et al.
Published: (2025)
by: Sun, Yifan, et al.
Published: (2025)
tinyBenchmarks: evaluating LLMs with fewer examples
by: Polo, Felipe Maia, et al.
Published: (2024)
by: Polo, Felipe Maia, et al.
Published: (2024)
BioDSA-1K: Benchmarking Data Science Agents for Biomedical Research
by: Wang, Zifeng, et al.
Published: (2025)
by: Wang, Zifeng, et al.
Published: (2025)
LABBench2: An Improved Benchmark for AI Systems Performing Biology Research
by: Laurent, Jon M, et al.
Published: (2026)
by: Laurent, Jon M, et al.
Published: (2026)
Beyond Size: How Gradients Shape Pruning Decisions in Large Language Models
by: Das, Rocktim Jyoti, et al.
Published: (2023)
by: Das, Rocktim Jyoti, et al.
Published: (2023)
Benchmarking Benchmark Leakage in Large Language Models
by: Xu, Ruijie, et al.
Published: (2024)
by: Xu, Ruijie, et al.
Published: (2024)
Exploring the Generalization of Cancer Clinical Trial Eligibility Classifiers Across Diseases
by: Yang, Yumeng, et al.
Published: (2024)
by: Yang, Yumeng, et al.
Published: (2024)
Reddit-Impacts: A Named Entity Recognition Dataset for Analyzing Clinical and Social Effects of Substance Use Derived from Social Media
by: Ge, Yao, et al.
Published: (2024)
by: Ge, Yao, et al.
Published: (2024)
AI-assisted Protocol Information Extraction For Improved Accuracy and Efficiency in Clinical Trial Workflows
by: Babaeipour, Ramtin, et al.
Published: (2026)
by: Babaeipour, Ramtin, et al.
Published: (2026)
DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios
by: Huang, Yao, et al.
Published: (2025)
by: Huang, Yao, et al.
Published: (2025)
Matching Patients to Clinical Trials with Large Language Models
by: Jin, Qiao, et al.
Published: (2023)
by: Jin, Qiao, et al.
Published: (2023)
Early Linguistic Pattern of Anxiety from Social Media Using Interpretable Linguistic Features: A Multi-Faceted Validation Study with Author-Disjoint Evaluation
by: Utsa, Arnab Das
Published: (2026)
by: Utsa, Arnab Das
Published: (2026)
Similar Items
-
SECRET: Semi-supervised Clinical Trial Document Similarity Search
by: Das, Trisha, et al.
Published: (2025) -
TrialSynth: Generation of Synthetic Sequential Clinical Trial Data
by: Gao, Chufan, et al.
Published: (2024) -
SynRL: Aligning Synthetic Clinical Trial Data with Human-preferred Clinical Endpoints Using Reinforcement Learning
by: Das, Trisha, et al.
Published: (2024) -
Automatically Labeling Clinical Trial Outcomes: A Large-Scale Benchmark for Drug Development
by: Gao, Chufan, et al.
Published: (2024) -
Synthetic Patient-Physician Dialogue Generation from Clinical Notes Using LLM
by: Das, Trisha, et al.
Published: (2024)