Towards Trustworthy Dermatology MLLMs: A Benchmark and Multimodal Evaluator for Diagnostic Narratives
Fuente:
arXiv
Salvato in:
| Autori principali: | Shen, Yuhao, Qian, Jiahe, Zhang, Shuping, Chen, Zhangtianyi, Lu, Tao, Zhou, Juexiao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CoTBox-TTT: Grounding Medical VQA with Visual Chain-of-Thought Boxes During Test-time Training
di: Qian, Jiahe, et al.
Pubblicazione: (2025)
di: Qian, Jiahe, et al.
Pubblicazione: (2025)
Trustworthy and Fair SkinGPT-R1 for Democratizing Dermatological Reasoning across Diverse Ethnicities
di: Shen, Yuhao, et al.
Pubblicazione: (2025)
di: Shen, Yuhao, et al.
Pubblicazione: (2025)
SkinGPT-X: A Self-Evolving Collaborative Multi-Agent System for Transparent and Trustworthy Dermatological Diagnosis
di: Chen, Zhangtianyi, et al.
Pubblicazione: (2026)
di: Chen, Zhangtianyi, et al.
Pubblicazione: (2026)
SkinCaRe: A Multimodal Dermatology Dataset Annotated with Medical Caption and Chain-of-Thought Reasoning
di: Shen, Yuhao, et al.
Pubblicazione: (2024)
di: Shen, Yuhao, et al.
Pubblicazione: (2024)
MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation
di: Huang, Jinsheng, et al.
Pubblicazione: (2024)
di: Huang, Jinsheng, et al.
Pubblicazione: (2024)
Drawing the Line: Enhancing Trustworthiness of MLLMs Through the Power of Refusal
di: Wang, Yuhao, et al.
Pubblicazione: (2024)
di: Wang, Yuhao, et al.
Pubblicazione: (2024)
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs
di: Zhao, Jiahe, et al.
Pubblicazione: (2025)
di: Zhao, Jiahe, et al.
Pubblicazione: (2025)
Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding
di: Wang, Youze, et al.
Pubblicazione: (2025)
di: Wang, Youze, et al.
Pubblicazione: (2025)
Tell Me Where You Are: Multimodal LLMs Meet Place Recognition
di: Lyu, Zonglin, et al.
Pubblicazione: (2024)
di: Lyu, Zonglin, et al.
Pubblicazione: (2024)
From GPT-4 to Gemini and Beyond: Assessing the Landscape of MLLMs on Generalizability, Trustworthiness and Causality through Four Modalities
di: Lu, Chaochao, et al.
Pubblicazione: (2024)
di: Lu, Chaochao, et al.
Pubblicazione: (2024)
Narrative Action Evaluation with Prompt-Guided Multimodal Interaction
di: Zhang, Shiyi, et al.
Pubblicazione: (2024)
di: Zhang, Shiyi, et al.
Pubblicazione: (2024)
Skin-R1: Toward Trustworthy Clinical Reasoning for Dermatological Diagnosis
di: Liu, Zehao, et al.
Pubblicazione: (2025)
di: Liu, Zehao, et al.
Pubblicazione: (2025)
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
di: Ouyang, Kun, et al.
Pubblicazione: (2024)
di: Ouyang, Kun, et al.
Pubblicazione: (2024)
RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback
di: Yu, Tianyu, et al.
Pubblicazione: (2023)
di: Yu, Tianyu, et al.
Pubblicazione: (2023)
EventBench: Towards Comprehensive Benchmarking of Event-based MLLMs
di: Liu, Shaoyu, et al.
Pubblicazione: (2025)
di: Liu, Shaoyu, et al.
Pubblicazione: (2025)
FunBench: Benchmarking Fundus Reading Skills of MLLMs
di: Wei, Qijie, et al.
Pubblicazione: (2025)
di: Wei, Qijie, et al.
Pubblicazione: (2025)
HERM: Benchmarking and Enhancing Multimodal LLMs for Human-Centric Understanding
di: Li, Keliang, et al.
Pubblicazione: (2024)
di: Li, Keliang, et al.
Pubblicazione: (2024)
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark
di: Cheng, Ziming, et al.
Pubblicazione: (2025)
di: Cheng, Ziming, et al.
Pubblicazione: (2025)
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs
di: Li, Qi, et al.
Pubblicazione: (2026)
di: Li, Qi, et al.
Pubblicazione: (2026)
Attention Reallocation: Towards Zero-cost and Controllable Hallucination Mitigation of MLLMs
di: Tu, Chongjun, et al.
Pubblicazione: (2025)
di: Tu, Chongjun, et al.
Pubblicazione: (2025)
Benchmarking Large and Small MLLMs
di: Feng, Xuelu, et al.
Pubblicazione: (2025)
di: Feng, Xuelu, et al.
Pubblicazione: (2025)
FREAK: A Fine-grained Hallucination Evaluation Benchmark for Advanced MLLMs
di: Yin, Zhihan, et al.
Pubblicazione: (2026)
di: Yin, Zhihan, et al.
Pubblicazione: (2026)
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
di: Lu, Lidong, et al.
Pubblicazione: (2025)
di: Lu, Lidong, et al.
Pubblicazione: (2025)
Are Multimodal LLMs Ready for Clinical Dermatology? A Real-World Evaluation in Dermatology
di: Jiang, Roy, et al.
Pubblicazione: (2026)
di: Jiang, Roy, et al.
Pubblicazione: (2026)
Decompose and Leverage Preferences from Expert Models for Improving Trustworthiness of MLLMs
di: Cao, Rui, et al.
Pubblicazione: (2024)
di: Cao, Rui, et al.
Pubblicazione: (2024)
Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs
di: Jiang, Xueying, et al.
Pubblicazione: (2026)
di: Jiang, Xueying, et al.
Pubblicazione: (2026)
IF-Bench: Benchmarking and Enhancing MLLMs for Infrared Images with Generative Visual Prompting
di: Zhang, Tao, et al.
Pubblicazione: (2025)
di: Zhang, Tao, et al.
Pubblicazione: (2025)
MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs
di: Yuan, Jiakang, et al.
Pubblicazione: (2025)
di: Yuan, Jiakang, et al.
Pubblicazione: (2025)
Decoupled Competitive Framework for Semi-supervised Medical Image Segmentation
di: Chen, Jiahe, et al.
Pubblicazione: (2025)
di: Chen, Jiahe, et al.
Pubblicazione: (2025)
MokA: Multimodal Low-Rank Adaptation for MLLMs
di: Wei, Yake, et al.
Pubblicazione: (2025)
di: Wei, Yake, et al.
Pubblicazione: (2025)
THEMIS: Towards Holistic Evaluation of MLLMs for Scientific Paper Fraud Forensics
di: Ma, Tzu-Yen, et al.
Pubblicazione: (2026)
di: Ma, Tzu-Yen, et al.
Pubblicazione: (2026)
AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation
di: Wang, Junyang, et al.
Pubblicazione: (2023)
di: Wang, Junyang, et al.
Pubblicazione: (2023)
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
di: Hao, Yunzhuo, et al.
Pubblicazione: (2025)
di: Hao, Yunzhuo, et al.
Pubblicazione: (2025)
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders
di: Fang, Bo, et al.
Pubblicazione: (2025)
di: Fang, Bo, et al.
Pubblicazione: (2025)
ActFormer: Scalable Collaborative Perception via Active Queries
di: Huang, Suozhi, et al.
Pubblicazione: (2024)
di: Huang, Suozhi, et al.
Pubblicazione: (2024)
iDETEX: Empowering MLLMs for Intelligent DETailed EXplainable IQA
di: Zhao, Zhaoran, et al.
Pubblicazione: (2025)
di: Zhao, Zhaoran, et al.
Pubblicazione: (2025)
MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI
di: Yao, Huanjin, et al.
Pubblicazione: (2025)
di: Yao, Huanjin, et al.
Pubblicazione: (2025)
Towards Benchmarking and Evaluating Deepfake Detection
di: Lin, Chenhao, et al.
Pubblicazione: (2022)
di: Lin, Chenhao, et al.
Pubblicazione: (2022)
EditBoard: Towards a Comprehensive Evaluation Benchmark for Text-Based Video Editing Models
di: Chen, Yupeng, et al.
Pubblicazione: (2024)
di: Chen, Yupeng, et al.
Pubblicazione: (2024)
HIS-GPT: Towards 3D Human-In-Scene Multimodal Understanding
di: Zhao, Jiahe, et al.
Pubblicazione: (2025)
di: Zhao, Jiahe, et al.
Pubblicazione: (2025)
Documenti analoghi
-
CoTBox-TTT: Grounding Medical VQA with Visual Chain-of-Thought Boxes During Test-time Training
di: Qian, Jiahe, et al.
Pubblicazione: (2025) -
Trustworthy and Fair SkinGPT-R1 for Democratizing Dermatological Reasoning across Diverse Ethnicities
di: Shen, Yuhao, et al.
Pubblicazione: (2025) -
SkinGPT-X: A Self-Evolving Collaborative Multi-Agent System for Transparent and Trustworthy Dermatological Diagnosis
di: Chen, Zhangtianyi, et al.
Pubblicazione: (2026) -
SkinCaRe: A Multimodal Dermatology Dataset Annotated with Medical Caption and Chain-of-Thought Reasoning
di: Shen, Yuhao, et al.
Pubblicazione: (2024) -
MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation
di: Huang, Jinsheng, et al.
Pubblicazione: (2024)