Multimodal Large Language Models for Medical Report Generation via Customized Prompt Tuning
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Chunlei, Hou, Jingyang, Shi, Yilei, Hu, Jingliang, Zhu, Xiao Xiang, Mou, Lichao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scale-Aware Contrastive Reverse Distillation for Unsupervised Medical Anomaly Detection
by: Li, Chunlei, et al.
Published: (2025)
by: Li, Chunlei, et al.
Published: (2025)
Reducing Annotation Burden: Exploiting Image Knowledge for Few-Shot Medical Video Object Segmentation via Spatiotemporal Consistency Relearning
by: Zheng, Zixuan, et al.
Published: (2025)
by: Zheng, Zixuan, et al.
Published: (2025)
UniCrossAdapter: Multimodal Adaptation of CLIP for Radiology Report Generation
by: Chen, Yaxiong, et al.
Published: (2025)
by: Chen, Yaxiong, et al.
Published: (2025)
Learning Domain-Aware Task Prompt Representations for Multi-Domain All-in-One Image Restoration
by: Dong, Guanglu, et al.
Published: (2026)
by: Dong, Guanglu, et al.
Published: (2026)
High-Order Progressive Trajectory Matching for Medical Image Dataset Distillation
by: Dong, Le, et al.
Published: (2025)
by: Dong, Le, et al.
Published: (2025)
Rethinking Cell Counting Methods: Decoupling Counting and Localization
by: Zheng, Zixuan, et al.
Published: (2025)
by: Zheng, Zixuan, et al.
Published: (2025)
ProPL: Universal Semi-Supervised Ultrasound Image Segmentation via Prompt-Guided Pseudo-Labeling
by: Chen, Yaxiong, et al.
Published: (2025)
by: Chen, Yaxiong, et al.
Published: (2025)
Taming Stable Diffusion for Computed Tomography Blind Super-Resolution
by: Li, Chunlei, et al.
Published: (2025)
by: Li, Chunlei, et al.
Published: (2025)
One-Shot Medical Video Object Segmentation via Temporal Contrastive Memory Networks
by: Chen, Yaxiong, et al.
Published: (2025)
by: Chen, Yaxiong, et al.
Published: (2025)
Dual Distillation for Few-Shot Anomaly Detection
by: Dong, Le, et al.
Published: (2026)
by: Dong, Le, et al.
Published: (2026)
Ultrasound Image-to-Video Synthesis via Latent Dynamic Diffusion Models
by: Chen, Tingxiu, et al.
Published: (2025)
by: Chen, Tingxiu, et al.
Published: (2025)
On Revisiting Entropy for Identifying Mislabeled Images
by: Li, Chunlei, et al.
Published: (2026)
by: Li, Chunlei, et al.
Published: (2026)
CausalCLIPSeg: Unlocking CLIP's Potential in Referring Medical Image Segmentation with Causal Intervention
by: Chen, Yaxiong, et al.
Published: (2025)
by: Chen, Yaxiong, et al.
Published: (2025)
Striving for Simplicity: Simple Yet Effective Prior-Aware Pseudo-Labeling for Semi-Supervised Ultrasound Image Segmentation
by: Chen, Yaxiong, et al.
Published: (2025)
by: Chen, Yaxiong, et al.
Published: (2025)
ChatEarthNet: A Global-Scale Image-Text Dataset Empowering Vision-Language Geo-Foundation Models
by: Yuan, Zhenghang, et al.
Published: (2024)
by: Yuan, Zhenghang, et al.
Published: (2024)
Enhancing Monocular Height Estimation via Weak Supervision from Imperfect Labels
by: Chen, Sining, et al.
Published: (2025)
by: Chen, Sining, et al.
Published: (2025)
Q-PART: Quasi-Periodic Adaptive Regression with Test-time Training for Pediatric Left Ventricular Ejection Fraction Regression
by: Liu, Jie, et al.
Published: (2025)
by: Liu, Jie, et al.
Published: (2025)
RRSIS: Referring Remote Sensing Image Segmentation
by: Yuan, Zhenghang, et al.
Published: (2023)
by: Yuan, Zhenghang, et al.
Published: (2023)
Event-Customized Image Generation
by: Wang, Zhen, et al.
Published: (2024)
by: Wang, Zhen, et al.
Published: (2024)
LDP: Parameter-Efficient Fine-Tuning of Multimodal LLM for Medical Report Generation
by: Zhou, Tianyu, et al.
Published: (2025)
by: Zhou, Tianyu, et al.
Published: (2025)
SemPT: Semantic Prompt Tuning for Vision-Language Models
by: Shi, Xiao, et al.
Published: (2025)
by: Shi, Xiao, et al.
Published: (2025)
Global Collinearity-aware Polygonizer for Polygonal Building Mapping in Remote Sensing
by: Zhang, Fahong, et al.
Published: (2025)
by: Zhang, Fahong, et al.
Published: (2025)
Reconstructing Building Height from Spaceborne TomoSAR Point Clouds Using a Dual-Topology Network
by: Chen, Zhaiyu, et al.
Published: (2026)
by: Chen, Zhaiyu, et al.
Published: (2026)
TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones
by: Yuan, Zhengqing, et al.
Published: (2023)
by: Yuan, Zhengqing, et al.
Published: (2023)
Parameter-Efficient Fine-Tuning Medical Multimodal Large Language Models for Medical Visual Grounding
by: He, Jinlong, et al.
Published: (2024)
by: He, Jinlong, et al.
Published: (2024)
HED-UNet: Combined Segmentation and Edge Detection for Monitoring the Antarctic Coastline
by: Heidler, Konrad, et al.
Published: (2021)
by: Heidler, Konrad, et al.
Published: (2021)
A Deep Active Contour Model for Delineating Glacier Calving Fronts
by: Heidler, Konrad, et al.
Published: (2023)
by: Heidler, Konrad, et al.
Published: (2023)
GlobalBuildingAtlas: An Open Global and Complete Dataset of Building Polygons, Heights and LoD1 3D Models
by: Zhu, Xiao Xiang, et al.
Published: (2025)
by: Zhu, Xiao Xiang, et al.
Published: (2025)
Med-2E3: A 2D-Enhanced 3D Medical Multimodal Large Language Model
by: Shi, Yiming, et al.
Published: (2024)
by: Shi, Yiming, et al.
Published: (2024)
Adapting Vision-Language Models to Open Classes via Test-Time Prompt Tuning
by: Gao, Zhengqing, et al.
Published: (2024)
by: Gao, Zhengqing, et al.
Published: (2024)
Large-scale flood modeling and forecasting with FloodCast
by: Xu, Qingsong, et al.
Published: (2024)
by: Xu, Qingsong, et al.
Published: (2024)
Vision-Driven Prompt Optimization for Large Language Models in Multimodal Generative Tasks
by: Franklin, Leo, et al.
Published: (2025)
by: Franklin, Leo, et al.
Published: (2025)
MePT: Multi-Representation Guided Prompt Tuning for Vision-Language Model
by: Wang, Xinyang, et al.
Published: (2024)
by: Wang, Xinyang, et al.
Published: (2024)
FedMVP: Federated Multimodal Visual Prompt Tuning for Vision-Language Models
by: Singha, Mainak, et al.
Published: (2025)
by: Singha, Mainak, et al.
Published: (2025)
HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation
by: Hu, Teng, et al.
Published: (2025)
by: Hu, Teng, et al.
Published: (2025)
SAM 3D for 3D Object Reconstruction from Remote Sensing Images
by: Yao, Junsheng, et al.
Published: (2025)
by: Yao, Junsheng, et al.
Published: (2025)
Effective Black-Box Multi-Faceted Attacks Breach Vision Large Language Model Guardrails
by: Yang, Yijun, et al.
Published: (2025)
by: Yang, Yijun, et al.
Published: (2025)
EarthNets: Empowering AI in Earth Observation
by: Xiong, Zhitong, et al.
Published: (2022)
by: Xiong, Zhitong, et al.
Published: (2022)
PolyGNN: Polyhedron-based Graph Neural Network for 3D Building Reconstruction from Point Clouds
by: Chen, Zhaiyu, et al.
Published: (2023)
by: Chen, Zhaiyu, et al.
Published: (2023)
Multi-Modal and Multi-Resolution Data Fusion for High-Resolution Cloud Removal: A Novel Baseline and Benchmark
by: Xu, Fang, et al.
Published: (2023)
by: Xu, Fang, et al.
Published: (2023)
Similar Items
-
Scale-Aware Contrastive Reverse Distillation for Unsupervised Medical Anomaly Detection
by: Li, Chunlei, et al.
Published: (2025) -
Reducing Annotation Burden: Exploiting Image Knowledge for Few-Shot Medical Video Object Segmentation via Spatiotemporal Consistency Relearning
by: Zheng, Zixuan, et al.
Published: (2025) -
UniCrossAdapter: Multimodal Adaptation of CLIP for Radiology Report Generation
by: Chen, Yaxiong, et al.
Published: (2025) -
Learning Domain-Aware Task Prompt Representations for Multi-Domain All-in-One Image Restoration
by: Dong, Guanglu, et al.
Published: (2026) -
High-Order Progressive Trajectory Matching for Medical Image Dataset Distillation
by: Dong, Le, et al.
Published: (2025)