HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Junying, Gui, Chi, Ouyang, Ruyi, Gao, Anningzhe, Chen, Shunian, Chen, Guiming Hardy, Wang, Xidong, Zhang, Ruifei, Cai, Zhenyang, Ji, Ke, Yu, Guangjun, Wan, Xiang, Wang, Benyou |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HuatuoGPT-II, One-stage Training for Medical Adaption of LLMs
by: Chen, Junying, et al.
Published: (2023)
by: Chen, Junying, et al.
Published: (2023)
HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
by: Chen, Junying, et al.
Published: (2024)
by: Chen, Junying, et al.
Published: (2024)
CoD, Towards an Interpretable Medical Agent using Chain of Diagnosis
by: Chen, Junying, et al.
Published: (2024)
by: Chen, Junying, et al.
Published: (2024)
ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
by: Chen, Guiming Hardy, et al.
Published: (2024)
by: Chen, Guiming Hardy, et al.
Published: (2024)
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
by: Chen, Junying, et al.
Published: (2025)
by: Chen, Junying, et al.
Published: (2025)
Humans or LLMs as the Judge? A Study on Judgement Biases
by: Chen, Guiming Hardy, et al.
Published: (2024)
by: Chen, Guiming Hardy, et al.
Published: (2024)
MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos
by: Wang, Rongsheng, et al.
Published: (2025)
by: Wang, Rongsheng, et al.
Published: (2025)
LLMs for Doctors: Leveraging Medical LLMs to Assist Doctors, Not Replace Them
by: Xie, Wenya, et al.
Published: (2024)
by: Xie, Wenya, et al.
Published: (2024)
MileBench: Benchmarking MLLMs in Long Context
by: Song, Dingjie, et al.
Published: (2024)
by: Song, Dingjie, et al.
Published: (2024)
MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria
by: Ge, Wentao, et al.
Published: (2023)
by: Ge, Wentao, et al.
Published: (2023)
LLMs Could Autonomously Learn Without External Supervision
by: Ji, Ke, et al.
Published: (2024)
by: Ji, Ke, et al.
Published: (2024)
Enabling Doctor-Centric Medical AI with LLMs through Workflow-Aligned Tasks and Benchmarks
by: Xie, Wenya, et al.
Published: (2025)
by: Xie, Wenya, et al.
Published: (2025)
LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via a Hybrid Architecture
by: Wang, Xidong, et al.
Published: (2024)
by: Wang, Xidong, et al.
Published: (2024)
Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs
by: Song, Dingjie, et al.
Published: (2024)
by: Song, Dingjie, et al.
Published: (2024)
ShizhenGPT: Towards Multimodal LLMs for Traditional Chinese Medicine
by: Chen, Junying, et al.
Published: (2025)
by: Chen, Junying, et al.
Published: (2025)
Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging
by: Cai, Zhenyang, et al.
Published: (2024)
by: Cai, Zhenyang, et al.
Published: (2024)
CMB: A Comprehensive Medical Benchmark in Chinese
by: Wang, Xidong, et al.
Published: (2023)
by: Wang, Xidong, et al.
Published: (2023)
Apollo: A Lightweight Multilingual Medical LLM towards Democratizing Medical AI to 6B People
by: Wang, Xidong, et al.
Published: (2024)
by: Wang, Xidong, et al.
Published: (2024)
Efficiently Democratizing Medical LLMs for 50 Languages via a Mixture of Language Family Experts
by: Zheng, Guorui, et al.
Published: (2024)
by: Zheng, Guorui, et al.
Published: (2024)
RAG-Instruct: Boosting LLMs with Diverse Retrieval-Augmented Instructions
by: Liu, Wanlong, et al.
Published: (2024)
by: Liu, Wanlong, et al.
Published: (2024)
OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning
by: Yu, Fei, et al.
Published: (2023)
by: Yu, Fei, et al.
Published: (2023)
LLMs for Mathematical Modeling: Towards Bridging the Gap between Natural and Mathematical Languages
by: Huang, Xuhan, et al.
Published: (2024)
by: Huang, Xuhan, et al.
Published: (2024)
Both Text and Images Leaked! A Systematic Analysis of Data Contamination in Multimodal LLM
by: Song, Dingjie, et al.
Published: (2024)
by: Song, Dingjie, et al.
Published: (2024)
Do LLMs Triage Like Clinicians? A Dynamic Study of Outpatient Referral
by: Liu, Xiaoxiao, et al.
Published: (2025)
by: Liu, Xiaoxiao, et al.
Published: (2025)
MicroVerse: A Preliminary Exploration Toward a Micro-World Simulation
by: Wang, Rongsheng, et al.
Published: (2026)
by: Wang, Rongsheng, et al.
Published: (2026)
Beyond ViT Tokens: Masked-Diffusion Pretrained Convolutional Pathology Foundation Model for Cell-Level Dense Prediction
by: Chen, Weiming, et al.
Published: (2026)
by: Chen, Weiming, et al.
Published: (2026)
DentalGPT: Incentivizing Multimodal Complex Reasoning in Dentistry
by: Cai, Zhenyang, et al.
Published: (2025)
by: Cai, Zhenyang, et al.
Published: (2025)
WaveMind: Towards a Conversational EEG Foundation Model Aligned to Textual and Visual Modalities
by: Zeng, Ziyi, et al.
Published: (2025)
by: Zeng, Ziyi, et al.
Published: (2025)
Agentifying Patient Dynamics within LLMs through Interacting with Clinical World Model
by: Wu, Minghao, et al.
Published: (2026)
by: Wu, Minghao, et al.
Published: (2026)
PrinciplismQA: A Philosophy-Grounded Approach to Assessing LLM-Human Clinical Medical Ethics Alignment
by: Hong, Chang, et al.
Published: (2025)
by: Hong, Chang, et al.
Published: (2025)
Can Editing LLMs Inject Harm?
by: Chen, Canyu, et al.
Published: (2024)
by: Chen, Canyu, et al.
Published: (2024)
A Comprehensive Analysis for Visual Object Hallucination in Large Vision-Language Models
by: Jing, Liqiang, et al.
Published: (2025)
by: Jing, Liqiang, et al.
Published: (2025)
TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis
by: Chen, Shunian, et al.
Published: (2025)
by: Chen, Shunian, et al.
Published: (2025)
VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment
by: Li, Lei, et al.
Published: (2024)
by: Li, Lei, et al.
Published: (2024)
Large Multimodal Agents: A Survey
by: Xie, Junlin, et al.
Published: (2024)
by: Xie, Junlin, et al.
Published: (2024)
To What Extent Do Token-Level Representations from Pathology Foundation Models Improve Dense Prediction?
by: Chen, Weiming, et al.
Published: (2026)
by: Chen, Weiming, et al.
Published: (2026)
Towards Codable Watermarking for Injecting Multi-bits Information to LLMs
by: Wang, Lean, et al.
Published: (2023)
by: Wang, Lean, et al.
Published: (2023)
BlenderLLM: Training Large Language Models for Computer-Aided Design with Self-improvement
by: Du, Yuhao, et al.
Published: (2024)
by: Du, Yuhao, et al.
Published: (2024)
Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization
by: Lai, Zhengzhao, et al.
Published: (2025)
by: Lai, Zhengzhao, et al.
Published: (2025)
AceGPT, Localizing Large Language Models in Arabic
by: Huang, Huang, et al.
Published: (2023)
by: Huang, Huang, et al.
Published: (2023)
Similar Items
-
HuatuoGPT-II, One-stage Training for Medical Adaption of LLMs
by: Chen, Junying, et al.
Published: (2023) -
HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
by: Chen, Junying, et al.
Published: (2024) -
CoD, Towards an Interpretable Medical Agent using Chain of Diagnosis
by: Chen, Junying, et al.
Published: (2024) -
ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
by: Chen, Guiming Hardy, et al.
Published: (2024) -
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
by: Chen, Junying, et al.
Published: (2025)