Discrete Diffusion Models with MLLMs for Unified Medical Multimodal Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Mao, Jiawei, Wang, Yuhan, Chen, Lifeng, Zhao, Can, Tang, Yucheng, Yang, Dong, Qu, Liangqiong, Xu, Daguang, Zhou, Yuyin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MedSegFactory: Text-Guided Generation of Medical Image-Mask Pairs
di: Mao, Jiawei, et al.
Pubblicazione: (2025)
di: Mao, Jiawei, et al.
Pubblicazione: (2025)
Decoupled Residual Denoising Diffusion Models for Unified and Data Efficient Image-to-Image Translation
di: Lin, Ziyue, et al.
Pubblicazione: (2026)
di: Lin, Ziyue, et al.
Pubblicazione: (2026)
Unleashing the Potential of SAM for Medical Adaptation via Hierarchical Decoding
di: Cheng, Zhiheng, et al.
Pubblicazione: (2024)
di: Cheng, Zhiheng, et al.
Pubblicazione: (2024)
VISTA3D: A Unified Segmentation Foundation Model For 3D Medical Imaging
di: He, Yufan, et al.
Pubblicazione: (2024)
di: He, Yufan, et al.
Pubblicazione: (2024)
Residual Denoising Diffusion Models
di: Liu, Jiawei, et al.
Pubblicazione: (2023)
di: Liu, Jiawei, et al.
Pubblicazione: (2023)
MAISI-v2: Accelerated 3D High-Resolution Medical Image Synthesis with Rectified Flow and Region-specific Contrastive Loss
di: Zhao, Can, et al.
Pubblicazione: (2025)
di: Zhao, Can, et al.
Pubblicazione: (2025)
DVG-Diffusion: Dual-View Guided Diffusion Model for CT Reconstruction from X-Rays
di: Xie, Xing, et al.
Pubblicazione: (2025)
di: Xie, Xing, et al.
Pubblicazione: (2025)
FedVLMBench: Benchmarking Federated Fine-Tuning of Vision-Language Models
di: Zheng, Weiying, et al.
Pubblicazione: (2025)
di: Zheng, Weiying, et al.
Pubblicazione: (2025)
MedVLThinker: Simple Baselines for Multimodal Medical Reasoning
di: Huang, Xiaoke, et al.
Pubblicazione: (2025)
di: Huang, Xiaoke, et al.
Pubblicazione: (2025)
MAISI: Medical AI for Synthetic Imaging
di: Guo, Pengfei, et al.
Pubblicazione: (2024)
di: Guo, Pengfei, et al.
Pubblicazione: (2024)
A Short Review and Evaluation of SAM2's Performance in 3D CT Image Segmentation
di: He, Yufan, et al.
Pubblicazione: (2024)
di: He, Yufan, et al.
Pubblicazione: (2024)
Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion
di: Li, Lijiang, et al.
Pubblicazione: (2026)
di: Li, Lijiang, et al.
Pubblicazione: (2026)
Unleashing the Potential of Large Language Models for Text-to-Image Generation through Autoregressive Representation Alignment
di: Xie, Xing, et al.
Pubblicazione: (2025)
di: Xie, Xing, et al.
Pubblicazione: (2025)
Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens
di: Pan, Kaihang, et al.
Pubblicazione: (2025)
di: Pan, Kaihang, et al.
Pubblicazione: (2025)
MedUnifier: Unifying Vision-and-Language Pre-training on Medical Data with Vision Generation Task using Discrete Visual Representations
di: Zhang, Ziyang, et al.
Pubblicazione: (2025)
di: Zhang, Ziyang, et al.
Pubblicazione: (2025)
Medical Vision Generalist: Unifying Medical Imaging Tasks in Context
di: Ren, Sucheng, et al.
Pubblicazione: (2024)
di: Ren, Sucheng, et al.
Pubblicazione: (2024)
Unified Multimodal Discrete Diffusion
di: Swerdlow, Alexander, et al.
Pubblicazione: (2025)
di: Swerdlow, Alexander, et al.
Pubblicazione: (2025)
Restorer: Removing Multi-Degradation with All-Axis Attention and Prompt Guidance
di: Mao, Jiawei, et al.
Pubblicazione: (2024)
di: Mao, Jiawei, et al.
Pubblicazione: (2024)
GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset
di: Wang, Yuhan, et al.
Pubblicazione: (2025)
di: Wang, Yuhan, et al.
Pubblicazione: (2025)
UVE: Are MLLMs Unified Evaluators for AI-Generated Videos?
di: Liu, Yuanxin, et al.
Pubblicazione: (2025)
di: Liu, Yuanxin, et al.
Pubblicazione: (2025)
Customizing Visual Emotion Evaluation for MLLMs: An Open-vocabulary, Multifaceted, and Scalable Approach
di: Wu, Daiqing, et al.
Pubblicazione: (2025)
di: Wu, Daiqing, et al.
Pubblicazione: (2025)
TP-Seg: Task-Prototype Framework for Unified Medical Lesion Segmentation
di: Xu, Jiawei, et al.
Pubblicazione: (2026)
di: Xu, Jiawei, et al.
Pubblicazione: (2026)
VAP-Diffusion: Enriching Descriptions with MLLMs for Enhanced Medical Image Generation
di: Huang, Peng, et al.
Pubblicazione: (2025)
di: Huang, Peng, et al.
Pubblicazione: (2025)
Auto3DSeg for Brain Tumor Segmentation from 3D MRI in BraTS 2023 Challenge
di: Myronenko, Andriy, et al.
Pubblicazione: (2025)
di: Myronenko, Andriy, et al.
Pubblicazione: (2025)
Harnessing EHRs for Diffusion-based Anomaly Detection on Chest X-rays
di: Kim, Harim, et al.
Pubblicazione: (2025)
di: Kim, Harim, et al.
Pubblicazione: (2025)
Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
di: Chen, Jiayi, et al.
Pubblicazione: (2025)
di: Chen, Jiayi, et al.
Pubblicazione: (2025)
Fast-DDPM: Fast Denoising Diffusion Probabilistic Models for Medical Image-to-Image Generation
di: Jiang, Hongxu, et al.
Pubblicazione: (2024)
di: Jiang, Hongxu, et al.
Pubblicazione: (2024)
DoraCycle: Domain-Oriented Adaptation of Unified Generative Model in Multimodal Cycles
di: Zhao, Rui, et al.
Pubblicazione: (2025)
di: Zhao, Rui, et al.
Pubblicazione: (2025)
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
di: Hao, Yunzhuo, et al.
Pubblicazione: (2025)
di: Hao, Yunzhuo, et al.
Pubblicazione: (2025)
See Detail Say Clear: Towards Brain CT Report Generation via Pathological Clue-driven Representation Learning
di: Zheng, Chengxin, et al.
Pubblicazione: (2024)
di: Zheng, Chengxin, et al.
Pubblicazione: (2024)
A Unified and Controllable Framework for Layered Image Generation with Visual Effects
di: Yang, Jinrui, et al.
Pubblicazione: (2026)
di: Yang, Jinrui, et al.
Pubblicazione: (2026)
An Empirical Study on Configuring In-Context Learning Demonstrations for Unleashing MLLMs' Sentimental Perception Capability
di: Wu, Daiqing, et al.
Pubblicazione: (2025)
di: Wu, Daiqing, et al.
Pubblicazione: (2025)
SOWing Information: Cultivating Contextual Coherence with MLLMs in Image Generation
di: Pei, Yuhan, et al.
Pubblicazione: (2024)
di: Pei, Yuhan, et al.
Pubblicazione: (2024)
UNIMO-G: Unified Image Generation through Multimodal Conditional Diffusion
di: Li, Wei, et al.
Pubblicazione: (2024)
di: Li, Wei, et al.
Pubblicazione: (2024)
Text2CT: Towards 3D CT Volume Generation from Free-text Descriptions Using Diffusion Model
di: Guo, Pengfei, et al.
Pubblicazione: (2025)
di: Guo, Pengfei, et al.
Pubblicazione: (2025)
HSENet: Hybrid Spatial Encoding Network for 3D Medical Vision-Language Understanding
di: Shi, Yanzhao, et al.
Pubblicazione: (2025)
di: Shi, Yanzhao, et al.
Pubblicazione: (2025)
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs
di: Li, Qi, et al.
Pubblicazione: (2026)
di: Li, Qi, et al.
Pubblicazione: (2026)
FaceCat: Enhancing Face Recognition Security with a Unified Diffusion Model
di: Chen, Jiawei, et al.
Pubblicazione: (2024)
di: Chen, Jiawei, et al.
Pubblicazione: (2024)
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
di: Ouyang, Kun, et al.
Pubblicazione: (2024)
di: Ouyang, Kun, et al.
Pubblicazione: (2024)
Linking Perception, Confidence and Accuracy in MLLMs
di: Du, Yuetian, et al.
Pubblicazione: (2026)
di: Du, Yuetian, et al.
Pubblicazione: (2026)
Documenti analoghi
-
MedSegFactory: Text-Guided Generation of Medical Image-Mask Pairs
di: Mao, Jiawei, et al.
Pubblicazione: (2025) -
Decoupled Residual Denoising Diffusion Models for Unified and Data Efficient Image-to-Image Translation
di: Lin, Ziyue, et al.
Pubblicazione: (2026) -
Unleashing the Potential of SAM for Medical Adaptation via Hierarchical Decoding
di: Cheng, Zhiheng, et al.
Pubblicazione: (2024) -
VISTA3D: A Unified Segmentation Foundation Model For 3D Medical Imaging
di: He, Yufan, et al.
Pubblicazione: (2024) -
Residual Denoising Diffusion Models
di: Liu, Jiawei, et al.
Pubblicazione: (2023)