Scaling medical imaging report generation with multimodal reinforcement learning
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Qianchu, Zhang, Sheng, Qin, Guanghui, Gu, Yu, Jin, Ying, Preston, Sam, Xu, Yanbo, Kiblawi, Sid, Yim, Wen-wai, Ossowski, Tim, Naumann, Tristan, Wei, Mu, Poon, Hoifung |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains
di: Liu, Qianchu, et al.
Pubblicazione: (2025)
di: Liu, Qianchu, et al.
Pubblicazione: (2025)
Med-RLVR: Emerging Medical Reasoning from a 3B base model via reinforcement Learning
di: Zhang, Sheng, et al.
Pubblicazione: (2025)
di: Zhang, Sheng, et al.
Pubblicazione: (2025)
Boltzmann Attention Sampling for Image Analysis with Small Objects
di: Zhao, Theodore, et al.
Pubblicazione: (2025)
di: Zhao, Theodore, et al.
Pubblicazione: (2025)
OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning
di: Ossowski, Timothy, et al.
Pubblicazione: (2025)
di: Ossowski, Timothy, et al.
Pubblicazione: (2025)
Be My Eyes: Extending Large Language Models to New Modalities Through Multi-Agent Collaboration
di: Huang, James Y., et al.
Pubblicazione: (2025)
di: Huang, James Y., et al.
Pubblicazione: (2025)
Learning Sparse Visual Representations via Spatial-Semantic Factorization
di: Zhao, Theodore Zhengde, et al.
Pubblicazione: (2026)
di: Zhao, Theodore Zhengde, et al.
Pubblicazione: (2026)
Exploring Scaling Laws for EHR Foundation Models
di: Zhang, Sheng, et al.
Pubblicazione: (2025)
di: Zhang, Sheng, et al.
Pubblicazione: (2025)
Universal Abstraction: Harnessing Frontier Models to Structure Real-World Data at Scale
di: Wong, Cliff, et al.
Pubblicazione: (2025)
di: Wong, Cliff, et al.
Pubblicazione: (2025)
Pareto Optimal Learning for Estimating Large Language Model Errors
di: Zhao, Theodore, et al.
Pubblicazione: (2023)
di: Zhao, Theodore, et al.
Pubblicazione: (2023)
BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs
di: Zhang, Sheng, et al.
Pubblicazione: (2023)
di: Zhang, Sheng, et al.
Pubblicazione: (2023)
DocLens: Multi-aspect Fine-grained Evaluation for Medical Text Generation
di: Xie, Yiqing, et al.
Pubblicazione: (2023)
di: Xie, Yiqing, et al.
Pubblicazione: (2023)
UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity Recognition
di: Zhou, Wenxuan, et al.
Pubblicazione: (2023)
di: Zhou, Wenxuan, et al.
Pubblicazione: (2023)
Attribute Structuring Improves LLM-Based Evaluation of Clinical Text Summaries
di: Gero, Zelalem, et al.
Pubblicazione: (2024)
di: Gero, Zelalem, et al.
Pubblicazione: (2024)
From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning
di: Xu, Nan, et al.
Pubblicazione: (2024)
di: Xu, Nan, et al.
Pubblicazione: (2024)
BiomedParse: a biomedical foundation model for image parsing of everything everywhere all at once
di: Zhao, Theodore, et al.
Pubblicazione: (2024)
di: Zhao, Theodore, et al.
Pubblicazione: (2024)
ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation
di: Liu, Qin, et al.
Pubblicazione: (2025)
di: Liu, Qin, et al.
Pubblicazione: (2025)
T-Rex: Text-assisted Retrosynthesis Prediction
di: Liu, Yifeng, et al.
Pubblicazione: (2024)
di: Liu, Yifeng, et al.
Pubblicazione: (2024)
MedGEN-Bench: Contextually entangled benchmark for open-ended multimodal medical generation
di: Yang, Junjie, et al.
Pubblicazione: (2025)
di: Yang, Junjie, et al.
Pubblicazione: (2025)
mDPO: Conditional Preference Optimization for Multimodal Large Language Models
di: Wang, Fei, et al.
Pubblicazione: (2024)
di: Wang, Fei, et al.
Pubblicazione: (2024)
Foundation Models for Biomedical Image Segmentation: A Survey
di: Lee, Ho Hin, et al.
Pubblicazione: (2024)
di: Lee, Ho Hin, et al.
Pubblicazione: (2024)
MetaScale: Test-Time Scaling with Evolving Meta-Thoughts
di: Liu, Qin, et al.
Pubblicazione: (2025)
di: Liu, Qin, et al.
Pubblicazione: (2025)
OmniStruct: Universal Text-to-Structure Generation across Diverse Schemas
di: Huang, James Y., et al.
Pubblicazione: (2025)
di: Huang, James Y., et al.
Pubblicazione: (2025)
AURAD: Anatomy-Pathology Unified Radiology Synthesis with Progressive Representations
di: Ding, Shuhan, et al.
Pubblicazione: (2025)
di: Ding, Shuhan, et al.
Pubblicazione: (2025)
Semantic-Clipping: Efficient Vision-Language Modeling with Semantic-Guidedd Visual Selection
di: Li, Bangzheng, et al.
Pubblicazione: (2025)
di: Li, Bangzheng, et al.
Pubblicazione: (2025)
Towards a clinically accessible radiology foundation model: open-access and lightweight, with automated evaluation
di: Chaves, Juan Manuel Zambrano, et al.
Pubblicazione: (2024)
di: Chaves, Juan Manuel Zambrano, et al.
Pubblicazione: (2024)
MAIRA-1: A specialised large multimodal model for radiology report generation
di: Hyland, Stephanie L., et al.
Pubblicazione: (2023)
di: Hyland, Stephanie L., et al.
Pubblicazione: (2023)
Offset Unlearning for Large Language Models
di: Huang, James Y., et al.
Pubblicazione: (2024)
di: Huang, James Y., et al.
Pubblicazione: (2024)
The Illusion of Readiness in Health AI
di: Gu, Yu, et al.
Pubblicazione: (2025)
di: Gu, Yu, et al.
Pubblicazione: (2025)
Video Models Can Reason with Verifiable Rewards
di: Zhu, Tinghui, et al.
Pubblicazione: (2026)
di: Zhu, Tinghui, et al.
Pubblicazione: (2026)
DermaVQA-DAS: Dermatology Assessment Schema (DAS) & Datasets for Closed-Ended Question Answering & Segmentation in Patient-Generated Dermatology Images
di: Yim, Wen-wai, et al.
Pubblicazione: (2025)
di: Yim, Wen-wai, et al.
Pubblicazione: (2025)
OLIVE: Object Level In-Context Visual Embeddings
di: Ossowski, Timothy, et al.
Pubblicazione: (2024)
di: Ossowski, Timothy, et al.
Pubblicazione: (2024)
RADAR: A Multimodal Benchmark for 3D Image-Based Radiology Report Review
di: Sun, Zhaoyi, et al.
Pubblicazione: (2026)
di: Sun, Zhaoyi, et al.
Pubblicazione: (2026)
MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes
di: Abacha, Asma Ben, et al.
Pubblicazione: (2024)
di: Abacha, Asma Ben, et al.
Pubblicazione: (2024)
WorldMedQA-V: a multilingual, multimodal medical examination dataset for multimodal language models evaluation
di: Matos, João, et al.
Pubblicazione: (2024)
di: Matos, João, et al.
Pubblicazione: (2024)
Closing the gap in multimodal medical representation alignment
di: Grassucci, Eleonora, et al.
Pubblicazione: (2026)
di: Grassucci, Eleonora, et al.
Pubblicazione: (2026)
Generative Medical Event Models Improve with Scale
di: Waxler, Shane, et al.
Pubblicazione: (2025)
di: Waxler, Shane, et al.
Pubblicazione: (2025)
A mechanism-driven reinforcement learning framework for shape optimization of airfoils
di: Wang, Jingfeng, et al.
Pubblicazione: (2024)
di: Wang, Jingfeng, et al.
Pubblicazione: (2024)
LLM generated responses to mitigate the impact of hate speech
di: Podolak, Jakub, et al.
Pubblicazione: (2023)
di: Podolak, Jakub, et al.
Pubblicazione: (2023)
MedViLaM: A multimodal large language model with advanced generalizability and explainability for medical data understanding and generation
di: Xu, Lijian, et al.
Pubblicazione: (2024)
di: Xu, Lijian, et al.
Pubblicazione: (2024)
Chapter GenRecipe for Generating Recipes from Videos through Deep Learning
di: Sin-wai, Chan
Pubblicazione: (2026)
di: Sin-wai, Chan
Pubblicazione: (2026)
Documenti analoghi
-
X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains
di: Liu, Qianchu, et al.
Pubblicazione: (2025) -
Med-RLVR: Emerging Medical Reasoning from a 3B base model via reinforcement Learning
di: Zhang, Sheng, et al.
Pubblicazione: (2025) -
Boltzmann Attention Sampling for Image Analysis with Small Objects
di: Zhao, Theodore, et al.
Pubblicazione: (2025) -
OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning
di: Ossowski, Timothy, et al.
Pubblicazione: (2025) -
Be My Eyes: Extending Large Language Models to New Modalities Through Multi-Agent Collaboration
di: Huang, James Y., et al.
Pubblicazione: (2025)