From Compound Figures to Composite Understanding: Developing a Multi-Modal LLM from Biomedical Literature with Medical Multiple-Image Benchmarking and Validation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Zhen, Fu, Yihang, Madera, Gabriel, Giuffre, Mauro, Applebaum, Serina, Kim, Hyunjae, Xu, Hua, Chen, Qingyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Benchmarking Direct Preference Optimization for Medical Large Vision-Language Models
von: Kim, Dain, et al.
Veröffentlicht: (2026)
von: Kim, Dain, et al.
Veröffentlicht: (2026)
Augmenting Biomedical Named Entity Recognition with General-domain Resources
von: Yin, Yu, et al.
Veröffentlicht: (2024)
von: Yin, Yu, et al.
Veröffentlicht: (2024)
Humans and Large Language Models in Clinical Decision Support: A Study with Medical Calculators
von: Wan, Nicholas, et al.
Veröffentlicht: (2024)
von: Wan, Nicholas, et al.
Veröffentlicht: (2024)
A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language Models
von: Liu, Jie, et al.
Veröffentlicht: (2024)
von: Liu, Jie, et al.
Veröffentlicht: (2024)
BioXArena: Benchmarking LLM Agents on Multi-Modal Biomedical Machine Learning Tasks
von: Li, Loka, et al.
Veröffentlicht: (2026)
von: Li, Loka, et al.
Veröffentlicht: (2026)
ChatBench: From Static Benchmarks to Human-AI Evaluation
von: Chang, Serina, et al.
Veröffentlicht: (2025)
von: Chang, Serina, et al.
Veröffentlicht: (2025)
On the Wasserstein median of probability measures
von: You, Kisung, et al.
Veröffentlicht: (2022)
von: You, Kisung, et al.
Veröffentlicht: (2022)
LEME: Open Large Language Models for Ophthalmology with Advanced Reasoning and Clinical Validation
von: Kim, Hyunjae, et al.
Veröffentlicht: (2024)
von: Kim, Hyunjae, et al.
Veröffentlicht: (2024)
Learning from Negative Samples in Biomedical Generative Entity Linking
von: Kim, Chanhwi, et al.
Veröffentlicht: (2024)
von: Kim, Chanhwi, et al.
Veröffentlicht: (2024)
MedCalc-Bench: Evaluating Large Language Models for Medical Calculations
von: Khandekar, Nikhil, et al.
Veröffentlicht: (2024)
von: Khandekar, Nikhil, et al.
Veröffentlicht: (2024)
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
von: Zuo, Yuxin, et al.
Veröffentlicht: (2025)
von: Zuo, Yuxin, et al.
Veröffentlicht: (2025)
Med-PRM: Medical Reasoning Models with Stepwise, Guideline-verified Process Rewards
von: Yun, Jaehoon, et al.
Veröffentlicht: (2025)
von: Yun, Jaehoon, et al.
Veröffentlicht: (2025)
ctELM: Decoding and Manipulating Embeddings of Clinical Trials with Embedding Language Models
von: Ondov, Brian, et al.
Veröffentlicht: (2026)
von: Ondov, Brian, et al.
Veröffentlicht: (2026)
MedINST: Meta Dataset of Biomedical Instructions
von: Han, Wenhan, et al.
Veröffentlicht: (2024)
von: Han, Wenhan, et al.
Veröffentlicht: (2024)
PubTator 3.0: an AI-powered Literature Resource for Unlocking Biomedical Knowledge
von: Wei, Chih-Hsuan, et al.
Veröffentlicht: (2024)
von: Wei, Chih-Hsuan, et al.
Veröffentlicht: (2024)
Compound-QA: A Benchmark for Evaluating LLMs on Compound Questions
von: Hou, Yutao, et al.
Veröffentlicht: (2024)
von: Hou, Yutao, et al.
Veröffentlicht: (2024)
Quantifying the Utility of User Simulators for Building Collaborative LLM Assistants
von: Suh, Joseph, et al.
Veröffentlicht: (2026)
von: Suh, Joseph, et al.
Veröffentlicht: (2026)
GeneGPT: Augmenting Large Language Models with Domain Tools for Improved Access to Biomedical Information
von: Jin, Qiao, et al.
Veröffentlicht: (2023)
von: Jin, Qiao, et al.
Veröffentlicht: (2023)
Valid Survey Simulations with Limited Human Data: The Roles of Prompting, Fine-Tuning, and Rectification
von: Krsteski, Stefan, et al.
Veröffentlicht: (2025)
von: Krsteski, Stefan, et al.
Veröffentlicht: (2025)
Lagrange Duality and Compound Multi-Attention Transformer for Semi-Supervised Medical Image Segmentation
von: Zheng, Fuchen, et al.
Veröffentlicht: (2024)
von: Zheng, Fuchen, et al.
Veröffentlicht: (2024)
OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
von: Ye, Hanrong, et al.
Veröffentlicht: (2025)
von: Ye, Hanrong, et al.
Veröffentlicht: (2025)
Pub-Guard-LLM: Detecting Retracted Biomedical Articles with Reliable Explanations
von: Chen, Lihu, et al.
Veröffentlicht: (2025)
von: Chen, Lihu, et al.
Veröffentlicht: (2025)
BDIViz: An Interactive Visualization System for Biomedical Schema Matching with LLM-Powered Validation
von: Wu, Eden, et al.
Veröffentlicht: (2025)
von: Wu, Eden, et al.
Veröffentlicht: (2025)
Med-CoReasoner: Reducing Language Disparities in Medical Reasoning via Language-Informed Co-Reasoning
von: Gao, Fan, et al.
Veröffentlicht: (2026)
von: Gao, Fan, et al.
Veröffentlicht: (2026)
EHRNavigator: A Multi-Agent System for Patient-Level Clinical Question Answering over Heterogeneous Electronic Health Records
von: Qian, Lingfei, et al.
Veröffentlicht: (2026)
von: Qian, Lingfei, et al.
Veröffentlicht: (2026)
From Coarse to Fine: Benchmarking and Reward Modeling for Writing-Centric Generation Tasks
von: Ren, Qingyu, et al.
Veröffentlicht: (2026)
von: Ren, Qingyu, et al.
Veröffentlicht: (2026)
Benchmarking Egocentric Clinical Intent Understanding Capability for Medical Multimodal Large Language Models
von: Liu, Shaonan, et al.
Veröffentlicht: (2026)
von: Liu, Shaonan, et al.
Veröffentlicht: (2026)
SciAssess: Benchmarking LLM Proficiency in Scientific Literature Analysis
von: Cai, Hengxing, et al.
Veröffentlicht: (2024)
von: Cai, Hengxing, et al.
Veröffentlicht: (2024)
DrBenchmark: A Large Language Understanding Evaluation Benchmark for French Biomedical Domain
von: Labrak, Yanis, et al.
Veröffentlicht: (2024)
von: Labrak, Yanis, et al.
Veröffentlicht: (2024)
Augmenting Black-box LLMs with Medical Textbooks for Biomedical Question Answering
von: Wang, Yubo, et al.
Veröffentlicht: (2023)
von: Wang, Yubo, et al.
Veröffentlicht: (2023)
Fair Text to Medical Image Diffusion Model with Subgroup Distribution Aligned Tuning
von: Han, Xu, et al.
Veröffentlicht: (2024)
von: Han, Xu, et al.
Veröffentlicht: (2024)
Graph-Based Alternatives to LLMs for Human Simulation
von: Suh, Joseph, et al.
Veröffentlicht: (2025)
von: Suh, Joseph, et al.
Veröffentlicht: (2025)
Enhancing Cross-Modal Medical Image Segmentation through Compositionality
von: Eijpe, Aniek, et al.
Veröffentlicht: (2024)
von: Eijpe, Aniek, et al.
Veröffentlicht: (2024)
From Panel to Pixel: Zoom-In Vision-Language Pretraining from Biomedical Scientific Literature
von: Yuan, Kun, et al.
Veröffentlicht: (2025)
von: Yuan, Kun, et al.
Veröffentlicht: (2025)
Medical Large Language Model Benchmarks Should Prioritize Construct Validity
von: Alaa, Ahmed, et al.
Veröffentlicht: (2025)
von: Alaa, Ahmed, et al.
Veröffentlicht: (2025)
Enhancing Meme Emotion Understanding with Multi-Level Modality Enhancement and Dual-Stage Modal Fusion
von: Shi, Yi, et al.
Veröffentlicht: (2025)
von: Shi, Yi, et al.
Veröffentlicht: (2025)
Foundation Models to Unlock Real-World Evidence from Nationwide Medical Claims
von: Ma, Fan, et al.
Veröffentlicht: (2026)
von: Ma, Fan, et al.
Veröffentlicht: (2026)
Rethinking Retrieval-Augmented Generation for Medicine: A Large-Scale, Systematic Expert Evaluation and Practical Insights
von: Kim, Hyunjae, et al.
Veröffentlicht: (2025)
von: Kim, Hyunjae, et al.
Veröffentlicht: (2025)
Language Models are Surprisingly Fragile to Drug Names in Biomedical Benchmarks
von: Gallifant, Jack, et al.
Veröffentlicht: (2024)
von: Gallifant, Jack, et al.
Veröffentlicht: (2024)
Rationale-Guided Retrieval Augmented Generation for Medical Question Answering
von: Sohn, Jiwoong, et al.
Veröffentlicht: (2024)
von: Sohn, Jiwoong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Benchmarking Direct Preference Optimization for Medical Large Vision-Language Models
von: Kim, Dain, et al.
Veröffentlicht: (2026) -
Augmenting Biomedical Named Entity Recognition with General-domain Resources
von: Yin, Yu, et al.
Veröffentlicht: (2024) -
Humans and Large Language Models in Clinical Decision Support: A Study with Medical Calculators
von: Wan, Nicholas, et al.
Veröffentlicht: (2024) -
A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language Models
von: Liu, Jie, et al.
Veröffentlicht: (2024) -
BioXArena: Benchmarking LLM Agents on Multi-Modal Biomedical Machine Learning Tasks
von: Li, Loka, et al.
Veröffentlicht: (2026)