MedVista3D: Vision-Language Modeling for Reducing Diagnostic Errors in 3D CT Disease Detection, Understanding and Reporting
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yuheng, Chen, Yenho, Lai, Yuxiang, Zhong, Jike, Wildman, Vanessa, Yang, Xiaofeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Med-R1: Reinforcement Learning for Generalizable Medical Reasoning in Vision-Language Models
by: Lai, Yuxiang, et al.
Published: (2025)
by: Lai, Yuxiang, et al.
Published: (2025)
Patient-Specific Autoregressive Models for Organ Motion Prediction in Radiotherapy
by: Lai, Yuxiang, et al.
Published: (2025)
by: Lai, Yuxiang, et al.
Published: (2025)
Are Video Models Emerging as Zero-Shot Learners and Reasoners in Medical Imaging?
by: Lai, Yuxiang, et al.
Published: (2025)
by: Lai, Yuxiang, et al.
Published: (2025)
Context Matters: Learning Global Semantics via Object-Centric Representation
by: Zhong, Jike, et al.
Published: (2025)
by: Zhong, Jike, et al.
Published: (2025)
MedDINOv3: How to adapt vision foundation models for medical image segmentation?
by: Li, Yuheng, et al.
Published: (2025)
by: Li, Yuheng, et al.
Published: (2025)
RoMedFormer: A Rotary-Embedding Transformer Foundation Model for 3D Genito-Pelvic Structure Segmentation in MRI and CT
by: Li, Yuheng, et al.
Published: (2025)
by: Li, Yuheng, et al.
Published: (2025)
Universal CT Representations from Anatomy to Disease Phenotype through Agglomerative Pretraining
by: Li, Yuheng, et al.
Published: (2026)
by: Li, Yuheng, et al.
Published: (2026)
Detection and segmentation of brain metastases on MRI using 3D‐MedDCNet
by: Yizhou Wu, et al.
Published: (2025)
by: Yizhou Wu, et al.
Published: (2025)
Towards Universal Text-driven CT Image Segmentation
by: Li, Yuheng, et al.
Published: (2025)
by: Li, Yuheng, et al.
Published: (2025)
AnatoMask: Enhancing Medical Image Segmentation with Reconstruction-guided Self-masking
by: Li, Yuheng, et al.
Published: (2024)
by: Li, Yuheng, et al.
Published: (2024)
Vision-Language Models for Automated 3D PET/CT Report Generation
by: Jiao, Wenpei, et al.
Published: (2025)
by: Jiao, Wenpei, et al.
Published: (2025)
3D-CT-GPT: Generating 3D Radiology Reports through Integration of Large Vision-Language Models
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
MedRegion-CT: Region-Focused Multimodal LLM for Comprehensive 3D CT Report Generation
by: Kyung, Sunggu, et al.
Published: (2025)
by: Kyung, Sunggu, et al.
Published: (2025)
Med3D-R1: Incentivizing Clinical Reasoning in 3D Medical Vision-Language Models for Abnormality Diagnosis
by: Lai, Haoran, et al.
Published: (2026)
by: Lai, Haoran, et al.
Published: (2026)
Learning to Read Where to Look: Disease-Aware Vision-Language Pretraining for 3D CT
by: Ging, Simon, et al.
Published: (2026)
by: Ging, Simon, et al.
Published: (2026)
EEE-Bench: A Comprehensive Multimodal Electrical And Electronics Engineering Benchmark
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
Vista3D: Unravel the 3D Darkside of a Single Image
by: Shen, Qiuhong, et al.
Published: (2024)
by: Shen, Qiuhong, et al.
Published: (2024)
EchoVLM: Measurement-Grounded Multimodal Learning for Echocardiography
by: Li, Yuheng, et al.
Published: (2025)
by: Li, Yuheng, et al.
Published: (2025)
Unifying 2D and 3D Vision-Language Understanding
by: Jain, Ayush, et al.
Published: (2025)
by: Jain, Ayush, et al.
Published: (2025)
Med3DVLM: An Efficient Vision-Language Model for 3D Medical Image Analysis
by: Xin, Yu, et al.
Published: (2025)
by: Xin, Yu, et al.
Published: (2025)
MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports
by: Kyung, Sunggu, et al.
Published: (2025)
by: Kyung, Sunggu, et al.
Published: (2025)
MedLSAM: Localize and Segment Anything Model for 3D CT Images
by: Lei, Wenhui, et al.
Published: (2023)
by: Lei, Wenhui, et al.
Published: (2023)
MedPruner: Training-Free Hierarchical Token Pruning for Efficient 3D Medical Image Understanding in Vision-Language Models
by: Liu, Shengyuan, et al.
Published: (2026)
by: Liu, Shengyuan, et al.
Published: (2026)
Med3DInsight: Enhancing 3D Medical Image Understanding with 2D Multi-Modal Large Language Models
by: Chen, Qiuhui, et al.
Published: (2024)
by: Chen, Qiuhui, et al.
Published: (2024)
MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering
by: Xi, Suyang, et al.
Published: (2026)
by: Xi, Suyang, et al.
Published: (2026)
VisionGPT-3D: A Generalized Multimodal Agent for Enhanced 3D Vision Understanding
by: Kelly, Chris, et al.
Published: (2024)
by: Kelly, Chris, et al.
Published: (2024)
Can 3D Vision-Language Models Truly Understand Natural Language?
by: Deng, Weipeng, et al.
Published: (2024)
by: Deng, Weipeng, et al.
Published: (2024)
Clinical Priors Guided Lung Disease Detection in 3D CT Scans
by: Lu, Kejin, et al.
Published: (2026)
by: Lu, Kejin, et al.
Published: (2026)
Multi‐group deformable convolution network for 3D medical image segmentation
by: Yuheng Li, et al.
Published: (2025)
by: Yuheng Li, et al.
Published: (2025)
Lost in Volume: The CT-SpatialVQA Benchmark for Evaluating Semantic-Spatial Understanding of 3D Medical Vision-Language Models
by: Monon, Mashrafi, et al.
Published: (2026)
by: Monon, Mashrafi, et al.
Published: (2026)
3D-RAD: A Comprehensive 3D Radiology Med-VQA Dataset with Multi-Temporal Analysis and Diverse Diagnostic Tasks
by: Gai, Xiaotang, et al.
Published: (2025)
by: Gai, Xiaotang, et al.
Published: (2025)
LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
by: Man, Yunze, et al.
Published: (2025)
by: Man, Yunze, et al.
Published: (2025)
Cardiac-CLIP: A Vision-Language Foundation Model for 3D Cardiac CT Images
by: Hu, Yutao, et al.
Published: (2025)
by: Hu, Yutao, et al.
Published: (2025)
Generalizable 7T T1-map Synthesis from 1.5T and 3T T1 MRI with an Efficient Transformer Model
by: Eidex, Zach, et al.
Published: (2025)
by: Eidex, Zach, et al.
Published: (2025)
Urban Redevelopment and Modernity in Liverpool and Manchester, 1918-1939
by: Wildman, Charlotte
Published: (2022)
by: Wildman, Charlotte
Published: (2022)
Unifying 3D Vision-Language Understanding via Promptable Queries
by: Zhu, Ziyu, et al.
Published: (2024)
by: Zhu, Ziyu, et al.
Published: (2024)
Advancing Lung Disease Diagnosis in 3D CT Scans
by: Li, Qingqiu, et al.
Published: (2025)
by: Li, Qingqiu, et al.
Published: (2025)
MedSyn: Text-guided Anatomy-aware Synthesis of High-Fidelity 3D CT Images
by: Xu, Yanwu, et al.
Published: (2023)
by: Xu, Yanwu, et al.
Published: (2023)
Unsupervised Adaptation from FDG to PSMA PET/CT for 3D Lesion Detection under Label Shift
by: Liu, Xiaofeng, et al.
Published: (2026)
by: Liu, Xiaofeng, et al.
Published: (2026)
HMR3D: Hierarchical Multimodal Representation for 3D Scene Understanding with Large Vision-Language Model
by: Li, Chen, et al.
Published: (2025)
by: Li, Chen, et al.
Published: (2025)
Similar Items
-
Med-R1: Reinforcement Learning for Generalizable Medical Reasoning in Vision-Language Models
by: Lai, Yuxiang, et al.
Published: (2025) -
Patient-Specific Autoregressive Models for Organ Motion Prediction in Radiotherapy
by: Lai, Yuxiang, et al.
Published: (2025) -
Are Video Models Emerging as Zero-Shot Learners and Reasoners in Medical Imaging?
by: Lai, Yuxiang, et al.
Published: (2025) -
Context Matters: Learning Global Semantics via Object-Centric Representation
by: Zhong, Jike, et al.
Published: (2025) -
MedDINOv3: How to adapt vision foundation models for medical image segmentation?
by: Li, Yuheng, et al.
Published: (2025)