Benchmarking the Thinking Mode of Multimodal Large Language Models in Clinical Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hong, Jindong, Chen, Tianjie, Luo, Lingjie, Zheng, Chuanyang, Xu, Ting, Yu, Haibao, Qiu, Jianing, Chen, Qianzhong, Huang, Suning, Xu, Yan, Gui, Yong, He, Yijun, Sun, Jiankai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SynapseRoute: An Auto-Route Switching Framework on Dual-State Large Language Model
von: Zhang, Wencheng, et al.
Veröffentlicht: (2025)
von: Zhang, Wencheng, et al.
Veröffentlicht: (2025)
Diagnosing Shoulder Disorders Using Multimodal Large Language Models and Consumer-Grade Cameras
von: Hong, Jindong, et al.
Veröffentlicht: (2025)
von: Hong, Jindong, et al.
Veröffentlicht: (2025)
Aria-NeRF: Multimodal Egocentric View Synthesis
von: Sun, Jiankai, et al.
Veröffentlicht: (2023)
von: Sun, Jiankai, et al.
Veröffentlicht: (2023)
Breaking Lock-In: Preserving Steerability under Low-Data VLA Post-Training
von: Huang, Suning, et al.
Veröffentlicht: (2026)
von: Huang, Suning, et al.
Veröffentlicht: (2026)
ParticleFormer: A 3D Point Cloud World Model for Multi-Object, Multi-Material Robotic Manipulation
von: Huang, Suning, et al.
Veröffentlicht: (2025)
von: Huang, Suning, et al.
Veröffentlicht: (2025)
MEDSYN: Benchmarking Multi-EviDence SYNthesis in Complex Clinical Cases for Multimodal Large Language Models
von: Chen, Boqi, et al.
Veröffentlicht: (2026)
von: Chen, Boqi, et al.
Veröffentlicht: (2026)
Mask What Matters: Mitigating Object Hallucinations in Multimodal Large Language Models with Object-Aligned Visual Contrastive Decoding
von: Chen, Boqi, et al.
Veröffentlicht: (2026)
von: Chen, Boqi, et al.
Veröffentlicht: (2026)
GRaD-Nav++: Vision-Language Model Enabled Visual Drone Navigation with Gaussian Radiance Fields and Differentiable Dynamics
von: Chen, Qianzhong, et al.
Veröffentlicht: (2025)
von: Chen, Qianzhong, et al.
Veröffentlicht: (2025)
MM-Soc: Benchmarking Multimodal Large Language Models in Social Media Platforms
von: Jin, Yiqiao, et al.
Veröffentlicht: (2024)
von: Jin, Yiqiao, et al.
Veröffentlicht: (2024)
Evaluating Large Language Models on Spatial Tasks: A Multi-Task Benchmarking Study
von: Xu, Liuchang, et al.
Veröffentlicht: (2024)
von: Xu, Liuchang, et al.
Veröffentlicht: (2024)
See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
von: Wu, Zongru, et al.
Veröffentlicht: (2025)
von: Wu, Zongru, et al.
Veröffentlicht: (2025)
MMESGBench: Pioneering Multimodal Understanding and Complex Reasoning Benchmark for ESG Tasks
von: Zhang, Lei, et al.
Veröffentlicht: (2025)
von: Zhang, Lei, et al.
Veröffentlicht: (2025)
AI Hiring with LLMs: A Context-Aware and Explainable Multi-Agent Framework for Resume Screening
von: Lo, Frank P. -W., et al.
Veröffentlicht: (2025)
von: Lo, Frank P. -W., et al.
Veröffentlicht: (2025)
Benchmarking Egocentric Clinical Intent Understanding Capability for Medical Multimodal Large Language Models
von: Liu, Shaonan, et al.
Veröffentlicht: (2026)
von: Liu, Shaonan, et al.
Veröffentlicht: (2026)
"A good pun is its own reword": Can Large Language Models Understand Puns?
von: Xu, Zhijun, et al.
Veröffentlicht: (2024)
von: Xu, Zhijun, et al.
Veröffentlicht: (2024)
EEE-QA: Exploring Effective and Efficient Question-Answer Representations
von: Hu, Zhanghao, et al.
Veröffentlicht: (2024)
von: Hu, Zhanghao, et al.
Veröffentlicht: (2024)
Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models
von: Li, Yunxin, et al.
Veröffentlicht: (2025)
von: Li, Yunxin, et al.
Veröffentlicht: (2025)
Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts
von: Xu, Haolei, et al.
Veröffentlicht: (2026)
von: Xu, Haolei, et al.
Veröffentlicht: (2026)
The Linear Attention Resurrection in Vision Transformer
von: Zheng, Chuanyang
Veröffentlicht: (2025)
von: Zheng, Chuanyang
Veröffentlicht: (2025)
CMHG: A Dataset and Benchmark for Headline Generation of Minority Languages in China
von: Xu, Guixian, et al.
Veröffentlicht: (2025)
von: Xu, Guixian, et al.
Veröffentlicht: (2025)
PATS: Process-Level Adaptive Thinking Mode Switching
von: Wang, Yi, et al.
Veröffentlicht: (2025)
von: Wang, Yi, et al.
Veröffentlicht: (2025)
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
Marker or Markerless? Mode-Switchable Optical Tactile Sensing for Diverse Robot Tasks
von: Ou, Ni, et al.
Veröffentlicht: (2024)
von: Ou, Ni, et al.
Veröffentlicht: (2024)
Think in Safety: Unveiling and Mitigating Safety Alignment Collapse in Multimodal Large Reasoning Model
von: Lou, Xinyue, et al.
Veröffentlicht: (2025)
von: Lou, Xinyue, et al.
Veröffentlicht: (2025)
Human-Level and Beyond: Benchmarking Large Language Models Against Clinical Pharmacists in Prescription Review
von: Yang, Yan, et al.
Veröffentlicht: (2025)
von: Yang, Yan, et al.
Veröffentlicht: (2025)
Visual Question Decomposition on Multimodal Large Language Models
von: Zhang, Haowei, et al.
Veröffentlicht: (2024)
von: Zhang, Haowei, et al.
Veröffentlicht: (2024)
iFormer: Integrating ConvNet and Transformer for Mobile Application
von: Zheng, Chuanyang
Veröffentlicht: (2025)
von: Zheng, Chuanyang
Veröffentlicht: (2025)
When to Continue Thinking: Adaptive Thinking Mode Switching for Efficient Reasoning
von: Zhang, Xiaoyun, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaoyun, et al.
Veröffentlicht: (2025)
SeqPO-SiMT: Sequential Policy Optimization for Simultaneous Machine Translation
von: Xu, Ting, et al.
Veröffentlicht: (2025)
von: Xu, Ting, et al.
Veröffentlicht: (2025)
Benchmarking Gaslighting Negation Attacks Against Multimodal Large Language Models
von: Zhu, Bin, et al.
Veröffentlicht: (2025)
von: Zhu, Bin, et al.
Veröffentlicht: (2025)
From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
von: Zhu, Wenxin, et al.
Veröffentlicht: (2025)
von: Zhu, Wenxin, et al.
Veröffentlicht: (2025)
AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
von: Jiang, Xilin, et al.
Veröffentlicht: (2026)
von: Jiang, Xilin, et al.
Veröffentlicht: (2026)
Thinking Makes LLM Agents Introverted: How Mandatory Thinking Can Backfire in User-Engaged Agents
von: Li, Jiatong, et al.
Veröffentlicht: (2026)
von: Li, Jiatong, et al.
Veröffentlicht: (2026)
Thinking with Visual Abstract: Enhancing Multimodal Reasoning via Visual Abstraction
von: Liu, Dairu, et al.
Veröffentlicht: (2025)
von: Liu, Dairu, et al.
Veröffentlicht: (2025)
The Role of Smart Cities in Ethical Design Framework
von: Chen, Yijun
Veröffentlicht: (2025)
von: Chen, Yijun
Veröffentlicht: (2025)
Smoothing Grounding and Reasoning for MLLM-Powered GUI Agents with Query-Oriented Pivot Tasks
von: Wu, Zongru, et al.
Veröffentlicht: (2025)
von: Wu, Zongru, et al.
Veröffentlicht: (2025)
MM-IQ: Benchmarking Human-Like Abstraction and Reasoning in Multimodal Models
von: Cai, Huanqia, et al.
Veröffentlicht: (2025)
von: Cai, Huanqia, et al.
Veröffentlicht: (2025)
Your Agent is More Brittle Than You Think: Uncovering Indirect Injection Vulnerabilities in Agentic LLMs
von: Zhu, Wenhui, et al.
Veröffentlicht: (2026)
von: Zhu, Wenhui, et al.
Veröffentlicht: (2026)
TRIMS: Trajectory-Ranked Instruction Masked Supervision for Diffusion Language Models
von: Chen, Lingjie, et al.
Veröffentlicht: (2026)
von: Chen, Lingjie, et al.
Veröffentlicht: (2026)
UNIDOC-BENCH: A Unified Benchmark for Document-Centric Multimodal RAG
von: Peng, Xiangyu, et al.
Veröffentlicht: (2025)
von: Peng, Xiangyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SynapseRoute: An Auto-Route Switching Framework on Dual-State Large Language Model
von: Zhang, Wencheng, et al.
Veröffentlicht: (2025) -
Diagnosing Shoulder Disorders Using Multimodal Large Language Models and Consumer-Grade Cameras
von: Hong, Jindong, et al.
Veröffentlicht: (2025) -
Aria-NeRF: Multimodal Egocentric View Synthesis
von: Sun, Jiankai, et al.
Veröffentlicht: (2023) -
Breaking Lock-In: Preserving Steerability under Low-Data VLA Post-Training
von: Huang, Suning, et al.
Veröffentlicht: (2026) -
ParticleFormer: A 3D Point Cloud World Model for Multi-Object, Multi-Material Robotic Manipulation
von: Huang, Suning, et al.
Veröffentlicht: (2025)