Distribution-Based Masked Medical Vision-Language Model Using Structured Reports
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Gowda, Shreyank N, Zhang, Ruichi, Gu, Xiao, Weng, Ying, Yang, Lu |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Masks and Manuscripts: Advancing Medical Pre-training with End-to-End Masking and Narrative Structuring
par: Gowda, Shreyank N, et autres
Publié: (2024)
par: Gowda, Shreyank N, et autres
Publié: (2024)
Prototype-Enhanced Confidence Modeling for Cross-Modal Medical Image-Report Retrieval
par: Gowda, Shreyank N, et autres
Publié: (2025)
par: Gowda, Shreyank N, et autres
Publié: (2025)
Is Temporal Prompting All We Need For Limited Labeled Action Recognition?
par: Gowda, Shreyank N, et autres
Publié: (2025)
par: Gowda, Shreyank N, et autres
Publié: (2025)
CC-SAM: SAM with Cross-feature Attention and Context for Ultrasound Image Segmentation
par: Gowda, Shreyank N, et autres
Publié: (2024)
par: Gowda, Shreyank N, et autres
Publié: (2024)
Telling Stories for Common Sense Zero-Shot Action Recognition
par: Gowda, Shreyank N, et autres
Publié: (2023)
par: Gowda, Shreyank N, et autres
Publié: (2023)
CAPT: Class-Aware Prompt Tuning for Federated Long-Tailed Learning with Vision-Language Model
par: Hou, Shihao, et autres
Publié: (2025)
par: Hou, Shihao, et autres
Publié: (2025)
Reimagining Reality: A Comprehensive Survey of Video Inpainting Techniques
par: Gowda, Shreyank N, et autres
Publié: (2024)
par: Gowda, Shreyank N, et autres
Publié: (2024)
Anyone Can Jailbreak: Prompt-Based Attacks on LLMs and T2Is
par: Mustafa, Ahmed B, et autres
Publié: (2025)
par: Mustafa, Ahmed B, et autres
Publié: (2025)
Adversarial Augmentation Training Makes Action Recognition Models More Robust to Realistic Video Distribution Shifts
par: Kim, Kiyoon, et autres
Publié: (2024)
par: Kim, Kiyoon, et autres
Publié: (2024)
Compression as an Adversarial Amplifier Through Decision Space Reduction
par: Evans, Lewis, et autres
Publié: (2026)
par: Evans, Lewis, et autres
Publié: (2026)
FE-Adapter: Adapting Image-based Emotion Classifiers to Videos
par: Gowda, Shreyank N, et autres
Publié: (2024)
par: Gowda, Shreyank N, et autres
Publié: (2024)
Continual Learning Improves Zero-Shot Action Recognition
par: Gowda, Shreyank N, et autres
Publié: (2024)
par: Gowda, Shreyank N, et autres
Publié: (2024)
Adaptive Data Dropout: Towards Self-Regulated Learning in Deep Neural Networks
par: Gahir, Amar, et autres
Publié: (2026)
par: Gahir, Amar, et autres
Publié: (2026)
Low-Effort Jailbreak Attacks Against Text-to-Image Safety Filters
par: Mustafa, Ahmed B, et autres
Publié: (2026)
par: Mustafa, Ahmed B, et autres
Publié: (2026)
SECOS: Semantic Capture for Rigorous Classification in Open-World Semi-Supervised Learning
par: Liu, Hezhao, et autres
Publié: (2026)
par: Liu, Hezhao, et autres
Publié: (2026)
Twin Trigger Generative Networks for Backdoor Attacks against Object Detection
par: Li, Zhiying, et autres
Publié: (2024)
par: Li, Zhiying, et autres
Publié: (2024)
FATE: A Prompt-Tuning-Based Semi-Supervised Learning Framework for Extremely Limited Labeled Data
par: Liu, Hezhao, et autres
Publié: (2025)
par: Liu, Hezhao, et autres
Publié: (2025)
ZeroDiff++: Substantial Unseen Visual-semantic Correlation in Zero-shot Learning
par: Ye, Zihan, et autres
Publié: (2026)
par: Ye, Zihan, et autres
Publié: (2026)
Delving into Out-of-Distribution Detection with Medical Vision-Language Models
par: Ju, Lie, et autres
Publié: (2025)
par: Ju, Lie, et autres
Publié: (2025)
Interpretable Zero-shot Learning with Infinite Class Concepts
par: Ye, Zihan, et autres
Publié: (2025)
par: Ye, Zihan, et autres
Publié: (2025)
Adversarial Robustness in Zero-Shot Learning:An Empirical Study on Class and Concept-Level Vulnerabilities
par: Peng, Zhiyuan, et autres
Publié: (2025)
par: Peng, Zhiyuan, et autres
Publié: (2025)
Watt For What: Rethinking Deep Learning's Energy-Performance Relationship
par: Gowda, Shreyank N, et autres
Publié: (2023)
par: Gowda, Shreyank N, et autres
Publié: (2023)
Bridging the Projection Gap: Overcoming Projection Bias Through Parameterized Distance Learning
par: Zhang, Chong, et autres
Publié: (2023)
par: Zhang, Chong, et autres
Publié: (2023)
Robust 3D Brain MRI Inpainting with Random Masking Augmentation
par: Zhang, Juexin, et autres
Publié: (2025)
par: Zhang, Juexin, et autres
Publié: (2025)
MaskedCLIP: Bridging the Masked and CLIP Space for Semi-Supervised Medical Vision-Language Pre-training
par: Zhu, Lei, et autres
Publié: (2025)
par: Zhu, Lei, et autres
Publié: (2025)
MVT: Mask-Grounded Vision-Language Models for Taxonomy-Aligned Land-Cover Tagging
par: Chen, Siyi, et autres
Publié: (2025)
par: Chen, Siyi, et autres
Publié: (2025)
Performance is not All You Need: Sustainability Considerations for Algorithms
par: Li, Xiang, et autres
Publié: (2025)
par: Li, Xiang, et autres
Publié: (2025)
Long-Tailed Distribution-Aware Router For Mixture-of-Experts in Large Vision-Language Model
par: Cai, Chaoxiang, et autres
Publié: (2025)
par: Cai, Chaoxiang, et autres
Publié: (2025)
Detecting and Evaluating Medical Hallucinations in Large Vision Language Models
par: Chen, Jiawei, et autres
Publié: (2024)
par: Chen, Jiawei, et autres
Publié: (2024)
U-VLM: Hierarchical Vision Language Modeling for Report Generation
par: Shi, Pengcheng, et autres
Publié: (2026)
par: Shi, Pengcheng, et autres
Publié: (2026)
Progressive Data Dropout: An Embarrassingly Simple Approach to Faster Training
par: Sathiyanarayanan, Shriram M, et autres
Publié: (2025)
par: Sathiyanarayanan, Shriram M, et autres
Publié: (2025)
Masked Diffusion Vision-Language Models for Temporal Action Localization
par: Wang, Fengshun, et autres
Publié: (2026)
par: Wang, Fengshun, et autres
Publié: (2026)
Mask-Based Modeling for Neural Radiance Fields
par: Yang, Ganlin, et autres
Publié: (2023)
par: Yang, Ganlin, et autres
Publié: (2023)
Principles of Visual Tokens for Efficient Video Understanding
par: Hao, Xinyue, et autres
Publié: (2024)
par: Hao, Xinyue, et autres
Publié: (2024)
CUE: Concept-Aware Multi-Label Expansion to Mitigate Concept Confusion in Long-Tailed Learning
par: Zhang, Ruichi, et autres
Publié: (2026)
par: Zhang, Ruichi, et autres
Publié: (2026)
MedTri: A Platform for Structured Medical Report Normalization to Enhance Vision-Language Pretraining
par: Chu, Yuetan, et autres
Publié: (2026)
par: Chu, Yuetan, et autres
Publié: (2026)
Visual Alignment of Medical Vision-Language Models for Grounded Radiology Report Generation
par: Bose, Sarosij, et autres
Publié: (2025)
par: Bose, Sarosij, et autres
Publié: (2025)
Efficient Medical Vision-Language Alignment Through Adapting Masked Vision Models
par: Lian, Chenyu, et autres
Publié: (2025)
par: Lian, Chenyu, et autres
Publié: (2025)
From Text to Mask: Localizing Entities Using the Attention of Text-to-Image Diffusion Models
par: Xiao, Changming, et autres
Publié: (2023)
par: Xiao, Changming, et autres
Publié: (2023)
Road Rage Reasoning with Vision-language Models (VLMs): Task Definition and Evaluation Dataset
par: Weng, Yibing, et autres
Publié: (2025)
par: Weng, Yibing, et autres
Publié: (2025)
Documents similaires
-
Masks and Manuscripts: Advancing Medical Pre-training with End-to-End Masking and Narrative Structuring
par: Gowda, Shreyank N, et autres
Publié: (2024) -
Prototype-Enhanced Confidence Modeling for Cross-Modal Medical Image-Report Retrieval
par: Gowda, Shreyank N, et autres
Publié: (2025) -
Is Temporal Prompting All We Need For Limited Labeled Action Recognition?
par: Gowda, Shreyank N, et autres
Publié: (2025) -
CC-SAM: SAM with Cross-feature Attention and Context for Ultrasound Image Segmentation
par: Gowda, Shreyank N, et autres
Publié: (2024) -
Telling Stories for Common Sense Zero-Shot Action Recognition
par: Gowda, Shreyank N, et autres
Publié: (2023)