Can LLMs' Tuning Methods Work in Medical Multimodal Domain?
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Jiawei, Jiang, Yue, Yang, Dingkang, Li, Mingcheng, Wei, Jinjie, Qian, Ziyun, Zhang, Lihua |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficiency in Focus: LayerNorm as a Catalyst for Fine-tuning Medical Visual Language Pre-trained Models
by: Chen, Jiawei, et al.
Published: (2024)
by: Chen, Jiawei, et al.
Published: (2024)
MCCD: Multi-Agent Collaboration-based Compositional Diffusion for Complex Text-to-Image Generation
by: Li, Mingcheng, et al.
Published: (2025)
by: Li, Mingcheng, et al.
Published: (2025)
CoMT: Chain-of-Medical-Thought Reduces Hallucination in Medical Report Generation
by: Jiang, Yue, et al.
Published: (2024)
by: Jiang, Yue, et al.
Published: (2024)
BloomScene: Lightweight Structured 3D Gaussian Splatting for Crossmodal Scene Generation
by: Hou, Xiaolu, et al.
Published: (2025)
by: Hou, Xiaolu, et al.
Published: (2025)
Toward Robust Incomplete Multimodal Sentiment Analysis via Hierarchical Representation Learning
by: Li, Mingcheng, et al.
Published: (2024)
by: Li, Mingcheng, et al.
Published: (2024)
Faster Diffusion Action Segmentation
by: Wang, Shuaibing, et al.
Published: (2024)
by: Wang, Shuaibing, et al.
Published: (2024)
HandGCAT: Occlusion-Robust 3D Hand Mesh Reconstruction from Monocular Images
by: Wang, Shuaibing, et al.
Published: (2024)
by: Wang, Shuaibing, et al.
Published: (2024)
Detecting and Evaluating Medical Hallucinations in Large Vision Language Models
by: Chen, Jiawei, et al.
Published: (2024)
by: Chen, Jiawei, et al.
Published: (2024)
Skip and Skip: Segmenting Medical Images with Prompts
by: Chen, Jiawei, et al.
Published: (2024)
by: Chen, Jiawei, et al.
Published: (2024)
Correlation-Decoupled Knowledge Distillation for Multimodal Sentiment Analysis with Incomplete Modalities
by: Li, Mingcheng, et al.
Published: (2024)
by: Li, Mingcheng, et al.
Published: (2024)
PersonaAnimator: Personalized Motion Transfer from Unconstrained Videos
by: Qian, Ziyun, et al.
Published: (2025)
by: Qian, Ziyun, et al.
Published: (2025)
Improving Multimodal Sentiment Analysis via Modality Optimization and Dynamic Primary Modality Selection
by: Yang, Dingkang, et al.
Published: (2025)
by: Yang, Dingkang, et al.
Published: (2025)
SMCD: High Realism Motion Style Transfer via Mamba-based Diffusion
by: Qian, Ziyun, et al.
Published: (2024)
by: Qian, Ziyun, et al.
Published: (2024)
Asynchronous Multimodal Video Sequence Fusion via Learning Modality-Exclusive and -Agnostic Representations
by: Yang, Dingkang, et al.
Published: (2024)
by: Yang, Dingkang, et al.
Published: (2024)
Resolving Evidence Sparsity: Agentic Context Engineering for Long-Document Understanding
by: Liu, Keliang, et al.
Published: (2025)
by: Liu, Keliang, et al.
Published: (2025)
MISS: A Generative Pretraining and Finetuning Approach for Med-VQA
by: Chen, Jiawei, et al.
Published: (2024)
by: Chen, Jiawei, et al.
Published: (2024)
ProFocus: Proactive Perception and Focused Reasoning in Vision-and-Language Navigation
by: Xue, Wei, et al.
Published: (2026)
by: Xue, Wei, et al.
Published: (2026)
SatireDecoder: Visual Cascaded Decoupling for Enhancing Satirical Image Comprehension
by: Jiang, Yue, et al.
Published: (2025)
by: Jiang, Yue, et al.
Published: (2025)
FysicsWorld: A Unified Full-Modality Benchmark for Any-to-Any Understanding, Generation, and Reasoning
by: Jiang, Yue, et al.
Published: (2025)
by: Jiang, Yue, et al.
Published: (2025)
Robust Emotion Recognition in Context Debiasing
by: Yang, Dingkang, et al.
Published: (2024)
by: Yang, Dingkang, et al.
Published: (2024)
MaskBEV: Towards A Unified Framework for BEV Detection and Map Segmentation
by: Zhao, Xiao, et al.
Published: (2024)
by: Zhao, Xiao, et al.
Published: (2024)
Towards Multimodal Sentiment Analysis Debiasing via Bias Purification
by: Yang, Dingkang, et al.
Published: (2024)
by: Yang, Dingkang, et al.
Published: (2024)
OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization
by: Han, Minghao, et al.
Published: (2026)
by: Han, Minghao, et al.
Published: (2026)
HybridOcc: NeRF Enhanced Transformer-based Multi-Camera 3D Occupancy Prediction
by: Zhao, Xiao, et al.
Published: (2024)
by: Zhao, Xiao, et al.
Published: (2024)
MM-Snowball: Evaluating and Mitigating Hallucination Snowballing in Multimodal Multi-Turn Dialogue
by: Jiang, Yue, et al.
Published: (2026)
by: Jiang, Yue, et al.
Published: (2026)
Can We Evaluate Domain Adaptation Models Without Target-Domain Labels?
by: Yang, Jianfei, et al.
Published: (2023)
by: Yang, Jianfei, et al.
Published: (2023)
MSCPT: Few-shot Whole Slide Image Classification with Multi-scale and Context-focused Prompt Tuning
by: Han, Minghao, et al.
Published: (2024)
by: Han, Minghao, et al.
Published: (2024)
Adaptive Reinforcement for Open-ended Medical Reasoning via Semantic-Guided Reward Collapse Mitigation
by: Liu, Yizhou, et al.
Published: (2025)
by: Liu, Yizhou, et al.
Published: (2025)
Forging a Dynamic Memory: Retrieval-Guided Continual Learning for Generalist Medical Foundation Models
by: Chen, Zizhi, et al.
Published: (2025)
by: Chen, Zizhi, et al.
Published: (2025)
Large Vision-Language Models as Emotion Recognizers in Context Awareness
by: Lei, Yuxuan, et al.
Published: (2024)
by: Lei, Yuxuan, et al.
Published: (2024)
Parameter-Efficient Fine-Tuning in Spectral Domain for Point Cloud Learning
by: Liang, Dingkang, et al.
Published: (2024)
by: Liang, Dingkang, et al.
Published: (2024)
Towards Context-Aware Emotion Recognition Debiasing from a Causal Demystification Perspective via De-confounded Training
by: Yang, Dingkang, et al.
Published: (2024)
by: Yang, Dingkang, et al.
Published: (2024)
LDP: Parameter-Efficient Fine-Tuning of Multimodal LLM for Medical Report Generation
by: Zhou, Tianyu, et al.
Published: (2025)
by: Zhou, Tianyu, et al.
Published: (2025)
Instruction Tuning-free Visual Token Complement for Multimodal LLMs
by: Wang, Dongsheng, et al.
Published: (2024)
by: Wang, Dongsheng, et al.
Published: (2024)
NAUTILUS: A Large Multimodal Model for Underwater Scene Understanding
by: Xu, Wei, et al.
Published: (2025)
by: Xu, Wei, et al.
Published: (2025)
Multimodal Causal-Driven Representation Learning for Generalizable Medical Image Segmentation
by: Liang, Xusheng, et al.
Published: (2025)
by: Liang, Xusheng, et al.
Published: (2025)
Multimodal Cancer Survival Analysis via Hypergraph Learning with Cross-Modality Rebalance
by: Qu, Mingcheng, et al.
Published: (2025)
by: Qu, Mingcheng, et al.
Published: (2025)
De-confounded Data-free Knowledge Distillation for Handling Distribution Shifts
by: Wang, Yuzheng, et al.
Published: (2024)
by: Wang, Yuzheng, et al.
Published: (2024)
Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
by: Guan, Yiran, et al.
Published: (2026)
by: Guan, Yiran, et al.
Published: (2026)
Fusing Pixels and Genes: Spatially-Aware Learning in Computational Pathology
by: Han, Minghao, et al.
Published: (2026)
by: Han, Minghao, et al.
Published: (2026)
Similar Items
-
Efficiency in Focus: LayerNorm as a Catalyst for Fine-tuning Medical Visual Language Pre-trained Models
by: Chen, Jiawei, et al.
Published: (2024) -
MCCD: Multi-Agent Collaboration-based Compositional Diffusion for Complex Text-to-Image Generation
by: Li, Mingcheng, et al.
Published: (2025) -
CoMT: Chain-of-Medical-Thought Reduces Hallucination in Medical Report Generation
by: Jiang, Yue, et al.
Published: (2024) -
BloomScene: Lightweight Structured 3D Gaussian Splatting for Crossmodal Scene Generation
by: Hou, Xiaolu, et al.
Published: (2025) -
Toward Robust Incomplete Multimodal Sentiment Analysis via Hierarchical Representation Learning
by: Li, Mingcheng, et al.
Published: (2024)