Enhancing the vision-language foundation model with key semantic knowledge-emphasized report refinement
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Weijian, Li, Cheng, Yang, Hao, Liu, Jiarun, Liang, Yong, Zheng, Hairong, Wang, Shanshan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A multi-modal vision-language model for generalizable annotation-free pathology localization
von: Yang, Hao, et al.
Veröffentlicht: (2024)
von: Yang, Hao, et al.
Veröffentlicht: (2024)
Enhancing Representation in Medical Vision-Language Foundation Models via Multi-Scale Information Extraction Techniques
von: Huang, Weijian, et al.
Veröffentlicht: (2024)
von: Huang, Weijian, et al.
Veröffentlicht: (2024)
Enhancing Representation in Radiography-Reports Foundation Model: A Granular Alignment Algorithm Using Masked Contrastive Learning
von: Huang, Weijian, et al.
Veröffentlicht: (2023)
von: Huang, Weijian, et al.
Veröffentlicht: (2023)
Multimodal self-supervised learning for lesion localization
von: Yang, Hao, et al.
Veröffentlicht: (2024)
von: Yang, Hao, et al.
Veröffentlicht: (2024)
Bias-constrained multimodal intelligence for equitable and reliable clinical AI
von: Li, Cheng, et al.
Veröffentlicht: (2026)
von: Li, Cheng, et al.
Veröffentlicht: (2026)
MLIP: Medical Language-Image Pre-training with Masked Local Representation Learning
von: Liu, Jiarun, et al.
Veröffentlicht: (2024)
von: Liu, Jiarun, et al.
Veröffentlicht: (2024)
BioVFM-21M: Benchmarking and Scaling Self-Supervised Vision Foundation Models for Biomedical Image Analysis
von: Liu, Jiarun, et al.
Veröffentlicht: (2025)
von: Liu, Jiarun, et al.
Veröffentlicht: (2025)
Thinker: A vision-language foundation model for embodied intelligence
von: Pan, Baiyu, et al.
Veröffentlicht: (2026)
von: Pan, Baiyu, et al.
Veröffentlicht: (2026)
Expert-level vision-language foundation model for real-world radiology and comprehensive evaluation
von: Liu, Xiaohong, et al.
Veröffentlicht: (2024)
von: Liu, Xiaohong, et al.
Veröffentlicht: (2024)
Optimized Vessel Segmentation: A Structure-Agnostic Approach with Small Vessel Enhancement and Morphological Correction
von: Song, Dongning, et al.
Veröffentlicht: (2024)
von: Song, Dongning, et al.
Veröffentlicht: (2024)
Zero-shot segmentation of skin tumors in whole-slide images with vision-language foundation models
von: Moreno, Santiago, et al.
Veröffentlicht: (2025)
von: Moreno, Santiago, et al.
Veröffentlicht: (2025)
An integrated language-vision foundation model for conversational diagnostics and triaging in primary eye care
von: Da Soh, Zhi, et al.
Veröffentlicht: (2025)
von: Da Soh, Zhi, et al.
Veröffentlicht: (2025)
Novel class discovery meets foundation models for 3D semantic segmentation
von: Riz, Luigi, et al.
Veröffentlicht: (2023)
von: Riz, Luigi, et al.
Veröffentlicht: (2023)
A multimodal vision foundation model for generalizable knee pathology
von: Yu, Kang, et al.
Veröffentlicht: (2026)
von: Yu, Kang, et al.
Veröffentlicht: (2026)
Swin-UMamba: Mamba-based UNet with ImageNet-based pretraining
von: Liu, Jiarun, et al.
Veröffentlicht: (2024)
von: Liu, Jiarun, et al.
Veröffentlicht: (2024)
Are foundation models for computer vision good conformal predictors?
von: Fillioux, Leo, et al.
Veröffentlicht: (2024)
von: Fillioux, Leo, et al.
Veröffentlicht: (2024)
Combining inherent knowledge of vision-language models with unsupervised domain adaptation through strong-weak guidance
von: Westfechtel, Thomas, et al.
Veröffentlicht: (2023)
von: Westfechtel, Thomas, et al.
Veröffentlicht: (2023)
MedDINOv3: How to adapt vision foundation models for medical image segmentation?
von: Li, Yuheng, et al.
Veröffentlicht: (2025)
von: Li, Yuheng, et al.
Veröffentlicht: (2025)
VFM-SDM: A vision foundation model-based framework for training-free, marker-free, and calibration-free structural displacement measurement
von: Xian, Qingyu, et al.
Veröffentlicht: (2026)
von: Xian, Qingyu, et al.
Veröffentlicht: (2026)
Fine-tuning vision foundation model for crack segmentation in civil infrastructures
von: Ge, Kang, et al.
Veröffentlicht: (2023)
von: Ge, Kang, et al.
Veröffentlicht: (2023)
A feature refinement module for light-weight semantic segmentation network
von: Wang, Zhiyan, et al.
Veröffentlicht: (2024)
von: Wang, Zhiyan, et al.
Veröffentlicht: (2024)
ZeroDiff++: Substantial Unseen Visual-semantic Correlation in Zero-shot Learning
von: Ye, Zihan, et al.
Veröffentlicht: (2026)
von: Ye, Zihan, et al.
Veröffentlicht: (2026)
FoMo4Wheat: Toward reliable crop vision foundation models with globally curated data
von: Han, Bing, et al.
Veröffentlicht: (2025)
von: Han, Bing, et al.
Veröffentlicht: (2025)
An analysis of vision-language models for fabric retrieval
von: Giuliari, Francesco, et al.
Veröffentlicht: (2025)
von: Giuliari, Francesco, et al.
Veröffentlicht: (2025)
DeViDe: Faceted medical knowledge for improved medical vision-language pre-training
von: Luo, Haozhe, et al.
Veröffentlicht: (2024)
von: Luo, Haozhe, et al.
Veröffentlicht: (2024)
EXACT: an explainable anomaly-aware vision foundation model for analysis of 3D chest CT
von: Bai, Xuguang, et al.
Veröffentlicht: (2026)
von: Bai, Xuguang, et al.
Veröffentlicht: (2026)
Enhancing medical vision-language contrastive learning via inter-matching relation modelling
von: Li, Mingjian, et al.
Veröffentlicht: (2024)
von: Li, Mingjian, et al.
Veröffentlicht: (2024)
Teaching pathology foundation models to accurately predict gene expression with parameter efficient knowledge transfer
von: Pan, Shi, et al.
Veröffentlicht: (2025)
von: Pan, Shi, et al.
Veröffentlicht: (2025)
EyeCLIP: A visual-language foundation model for multi-modal ophthalmic image analysis
von: Shi, Danli, et al.
Veröffentlicht: (2024)
von: Shi, Danli, et al.
Veröffentlicht: (2024)
PB-IAD: Utilizing multimodal foundation models for semantic industrial anomaly detection in dynamic manufacturing environments
von: Hofmann, Bernd, et al.
Veröffentlicht: (2025)
von: Hofmann, Bernd, et al.
Veröffentlicht: (2025)
Improving vision-language alignment with graph spiking hybrid Networks
von: Zhang, Siyu, et al.
Veröffentlicht: (2025)
von: Zhang, Siyu, et al.
Veröffentlicht: (2025)
Visual hallucination detection in large vision-language models via evidential conflict
von: Huang, Tao, et al.
Veröffentlicht: (2025)
von: Huang, Tao, et al.
Veröffentlicht: (2025)
Object segmentation in the wild with foundation models: application to vision assisted neuro-prostheses for upper limbs
von: Atoki, Bolutife, et al.
Veröffentlicht: (2025)
von: Atoki, Bolutife, et al.
Veröffentlicht: (2025)
Attend what matters: Leveraging vision foundational models for breast cancer classification using mammograms
von: Sanghvi, Samyak, et al.
Veröffentlicht: (2026)
von: Sanghvi, Samyak, et al.
Veröffentlicht: (2026)
ProFound: A moderate-sized vision foundation model for multi-task prostate imaging
von: Wang, Yipei, et al.
Veröffentlicht: (2026)
von: Wang, Yipei, et al.
Veröffentlicht: (2026)
Do computer vision foundation models learn the low-level characteristics of the human visual system?
von: Cai, Yancheng, et al.
Veröffentlicht: (2025)
von: Cai, Yancheng, et al.
Veröffentlicht: (2025)
SamRobNODDI: Q-Space Sampling-Augmented Continuous Representation Learning for Robust and Generalized NODDI
von: Xiao, Taohui, et al.
Veröffentlicht: (2024)
von: Xiao, Taohui, et al.
Veröffentlicht: (2024)
Rapidly deploying on-device eye tracking by distilling visual foundation models
von: Jiang, Cheng, et al.
Veröffentlicht: (2026)
von: Jiang, Cheng, et al.
Veröffentlicht: (2026)
The Solution for the CVPR 2023 1st foundation model challenge-Track2
von: Xu, Haonan, et al.
Veröffentlicht: (2024)
von: Xu, Haonan, et al.
Veröffentlicht: (2024)
An efficient framework based on large foundation model for cervical cytopathology whole slide image screening
von: Huang, Jialong, et al.
Veröffentlicht: (2024)
von: Huang, Jialong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A multi-modal vision-language model for generalizable annotation-free pathology localization
von: Yang, Hao, et al.
Veröffentlicht: (2024) -
Enhancing Representation in Medical Vision-Language Foundation Models via Multi-Scale Information Extraction Techniques
von: Huang, Weijian, et al.
Veröffentlicht: (2024) -
Enhancing Representation in Radiography-Reports Foundation Model: A Granular Alignment Algorithm Using Masked Contrastive Learning
von: Huang, Weijian, et al.
Veröffentlicht: (2023) -
Multimodal self-supervised learning for lesion localization
von: Yang, Hao, et al.
Veröffentlicht: (2024) -
Bias-constrained multimodal intelligence for equitable and reliable clinical AI
von: Li, Cheng, et al.
Veröffentlicht: (2026)