MM-Skin: Enhancing Dermatology Vision-Language Model with an Image-Text Dataset Derived from Textbooks
Fuente:
arXiv
Salvato in:
| Autori principali: | Zeng, Wenqi, Sun, Yuqi, Ma, Chenxi, Tan, Weimin, Yan, Bo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MicroVQA++: High-Quality Microscopy Reasoning Dataset with Weakly Supervised Graphs for Multimodal Large Language Model
di: Li, Manyu, et al.
Pubblicazione: (2025)
di: Li, Manyu, et al.
Pubblicazione: (2025)
Unifying Segment Anything in Microscopy with Vision-Language Knowledge
di: Li, Manyu, et al.
Pubblicazione: (2025)
di: Li, Manyu, et al.
Pubblicazione: (2025)
A Medical Data-Effective Learning Benchmark for Highly Efficient Pre-training of Foundation Models
di: Yang, Wenxuan, et al.
Pubblicazione: (2024)
di: Yang, Wenxuan, et al.
Pubblicazione: (2024)
MicroWorld: Empowering Multimodal Large Language Models to Bridge the Microscopic Domain Gap with Multimodal Attribute Graph
di: Li, Manyu, et al.
Pubblicazione: (2026)
di: Li, Manyu, et al.
Pubblicazione: (2026)
Vision-Language Models as Differentiable Semantic and Spatial Rewards for Text-to-3D Generation
di: Bai, Weimin, et al.
Pubblicazione: (2025)
di: Bai, Weimin, et al.
Pubblicazione: (2025)
SkinGEN: an Explainable Dermatology Diagnosis-to-Generation Framework with Interactive Vision-Language Models
di: Lin, Bo, et al.
Pubblicazione: (2024)
di: Lin, Bo, et al.
Pubblicazione: (2024)
SkinCaRe: A Multimodal Dermatology Dataset Annotated with Medical Caption and Chain-of-Thought Reasoning
di: Shen, Yuhao, et al.
Pubblicazione: (2024)
di: Shen, Yuhao, et al.
Pubblicazione: (2024)
MM-ISTS: Cooperating Irregularly Sampled Time Series Forecasting with Multimodal Vision-Text LLMs
di: Lei, Zhi, et al.
Pubblicazione: (2026)
di: Lei, Zhi, et al.
Pubblicazione: (2026)
CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation
di: Chen, Wei, et al.
Pubblicazione: (2024)
di: Chen, Wei, et al.
Pubblicazione: (2024)
MAKE: Multi-Aspect Knowledge-Enhanced Vision-Language Pretraining for Zero-shot Dermatological Assessment
di: Yan, Siyuan, et al.
Pubblicazione: (2025)
di: Yan, Siyuan, et al.
Pubblicazione: (2025)
Derm1M: A Million-scale Vision-Language Dataset Aligned with Clinical Ontology Knowledge for Dermatology
di: Yan, Siyuan, et al.
Pubblicazione: (2025)
di: Yan, Siyuan, et al.
Pubblicazione: (2025)
MM-Retinal: Knowledge-Enhanced Foundational Pretraining with Fundus Image-Text Expertise
di: Wu, Ruiqi, et al.
Pubblicazione: (2024)
di: Wu, Ruiqi, et al.
Pubblicazione: (2024)
Enhanced Continual Learning of Vision-Language Models with Model Fusion
di: Gao, Haoyuan, et al.
Pubblicazione: (2025)
di: Gao, Haoyuan, et al.
Pubblicazione: (2025)
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models
di: Wei, Zhixiang, et al.
Pubblicazione: (2025)
di: Wei, Zhixiang, et al.
Pubblicazione: (2025)
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining
di: Zhang, Wenqi, et al.
Pubblicazione: (2025)
di: Zhang, Wenqi, et al.
Pubblicazione: (2025)
FacialFlowNet: Advancing Facial Optical Flow Estimation with a Diverse Dataset and a Decomposed Model
di: Lu, Jianzhi, et al.
Pubblicazione: (2024)
di: Lu, Jianzhi, et al.
Pubblicazione: (2024)
MM-R$^3$: On (In-)Consistency of Vision-Language Models (VLMs)
di: Chou, Shih-Han, et al.
Pubblicazione: (2024)
di: Chou, Shih-Han, et al.
Pubblicazione: (2024)
Enhancing Remote Sensing Vision-Language Models Through MLLM and LLM-Based High-Quality Image-Text Dataset Generation
di: He, Yiguo, et al.
Pubblicazione: (2025)
di: He, Yiguo, et al.
Pubblicazione: (2025)
TrojVLM: Backdoor Attack Against Vision Language Models
di: Lyu, Weimin, et al.
Pubblicazione: (2024)
di: Lyu, Weimin, et al.
Pubblicazione: (2024)
The Impact of Skin Tone Label Granularity on the Performance and Fairness of AI Based Dermatology Image Classification Models
di: Shah, Partha, et al.
Pubblicazione: (2025)
di: Shah, Partha, et al.
Pubblicazione: (2025)
EVLF: Early Vision-Language Fusion for Generative Dataset Distillation
di: Cai, Wenqi, et al.
Pubblicazione: (2026)
di: Cai, Wenqi, et al.
Pubblicazione: (2026)
CNNs, Transformers, Hybrid, and Vision Language Models for Skin Cancer Detection
di: Dey, Durjoy, et al.
Pubblicazione: (2026)
di: Dey, Durjoy, et al.
Pubblicazione: (2026)
Benchmarking Vision Language Model Unlearning via Fictitious Facial Identity Dataset
di: Ma, Yingzi, et al.
Pubblicazione: (2024)
di: Ma, Yingzi, et al.
Pubblicazione: (2024)
Trustworthy and Fair SkinGPT-R1 for Democratizing Dermatological Reasoning across Diverse Ethnicities
di: Shen, Yuhao, et al.
Pubblicazione: (2025)
di: Shen, Yuhao, et al.
Pubblicazione: (2025)
PASSION for Dermatology: Bridging the Diversity Gap with Pigmented Skin Images from Sub-Saharan Africa
di: Gottfrois, Philippe, et al.
Pubblicazione: (2024)
di: Gottfrois, Philippe, et al.
Pubblicazione: (2024)
A Generative Framework for Self-Supervised Facial Representation Learning
di: He, Ruian, et al.
Pubblicazione: (2023)
di: He, Ruian, et al.
Pubblicazione: (2023)
LAION-SG: An Enhanced Large-Scale Dataset for Training Complex Image-Text Models with Structural Annotations
di: Li, Zejian, et al.
Pubblicazione: (2024)
di: Li, Zejian, et al.
Pubblicazione: (2024)
Tag2Text: Guiding Vision-Language Model via Image Tagging
di: Huang, Xinyu, et al.
Pubblicazione: (2023)
di: Huang, Xinyu, et al.
Pubblicazione: (2023)
Addressing Imbalance for Class Incremental Learning in Medical Image Classification
di: Hao, Xuze, et al.
Pubblicazione: (2024)
di: Hao, Xuze, et al.
Pubblicazione: (2024)
A Multimodal Vision Foundation Model for Clinical Dermatology
di: Yan, Siyuan, et al.
Pubblicazione: (2024)
di: Yan, Siyuan, et al.
Pubblicazione: (2024)
Quality Text, Robust Vision: The Role of Language in Enhancing Visual Robustness of Vision-Language Models
di: Waseda, Futa, et al.
Pubblicazione: (2025)
di: Waseda, Futa, et al.
Pubblicazione: (2025)
Detecting Text Manipulation in Images using Vision Language Models
di: Vidit, Vidit, et al.
Pubblicazione: (2025)
di: Vidit, Vidit, et al.
Pubblicazione: (2025)
ChatEarthNet: A Global-Scale Image-Text Dataset Empowering Vision-Language Geo-Foundation Models
di: Yuan, Zhenghang, et al.
Pubblicazione: (2024)
di: Yuan, Zhenghang, et al.
Pubblicazione: (2024)
Understanding Degradation with Vision Language Model
di: Lan, Guanzhou, et al.
Pubblicazione: (2026)
di: Lan, Guanzhou, et al.
Pubblicazione: (2026)
DermaCon-IN: A Multi-concept Annotated Dermatological Image Dataset of Indian Skin Disorders for Clinical AI Research
di: Madarkar, Shanawaj S, et al.
Pubblicazione: (2025)
di: Madarkar, Shanawaj S, et al.
Pubblicazione: (2025)
Reading Images Like Texts: Sequential Image Understanding in Vision-Language Models
di: Li, Yueyan, et al.
Pubblicazione: (2025)
di: Li, Yueyan, et al.
Pubblicazione: (2025)
Enhancing Skin Lesion Diagnosis with Ensemble Learning
di: Liu, Xiaoyi, et al.
Pubblicazione: (2024)
di: Liu, Xiaoyi, et al.
Pubblicazione: (2024)
BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature
di: Lozano, Alejandro, et al.
Pubblicazione: (2025)
di: Lozano, Alejandro, et al.
Pubblicazione: (2025)
Toward Accessible Dermatology: Skin Lesion Classification Using Deep Learning Models on Mobile-Acquired Images
di: Newaz, Asif, et al.
Pubblicazione: (2025)
di: Newaz, Asif, et al.
Pubblicazione: (2025)
TransMed: Large Language Models Enhance Vision Transformer for Biomedical Image Classification
di: Zheng, Kaipeng, et al.
Pubblicazione: (2023)
di: Zheng, Kaipeng, et al.
Pubblicazione: (2023)
Documenti analoghi
-
MicroVQA++: High-Quality Microscopy Reasoning Dataset with Weakly Supervised Graphs for Multimodal Large Language Model
di: Li, Manyu, et al.
Pubblicazione: (2025) -
Unifying Segment Anything in Microscopy with Vision-Language Knowledge
di: Li, Manyu, et al.
Pubblicazione: (2025) -
A Medical Data-Effective Learning Benchmark for Highly Efficient Pre-training of Foundation Models
di: Yang, Wenxuan, et al.
Pubblicazione: (2024) -
MicroWorld: Empowering Multimodal Large Language Models to Bridge the Microscopic Domain Gap with Multimodal Attribute Graph
di: Li, Manyu, et al.
Pubblicazione: (2026) -
Vision-Language Models as Differentiable Semantic and Spatial Rewards for Text-to-3D Generation
di: Bai, Weimin, et al.
Pubblicazione: (2025)