Enhancing Abnormality Grounding for Vision Language Models with Knowledge Descriptions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Jun, Liu, Che, Bai, Wenjia, Arcucci, Rossella, Bercea, Cosmin I., Schnabel, Julia A. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Knowledge to Sight: Reasoning over Visual Attributes via Knowledge Decomposition for Abnormality Grounding
von: Li, Jun, et al.
Veröffentlicht: (2025)
von: Li, Jun, et al.
Veröffentlicht: (2025)
FMBench: Benchmarking Fairness in Multimodal Large Language Models on Medical Tasks
von: Wu, Peiran, et al.
Veröffentlicht: (2024)
von: Wu, Peiran, et al.
Veröffentlicht: (2024)
Utilizing Synthetic Data for Medical Vision-Language Pre-training: Bypassing the Need for Real Images
von: Liu, Che, et al.
Veröffentlicht: (2023)
von: Liu, Che, et al.
Veröffentlicht: (2023)
How Far Have Medical Vision-Language Models Come? A Comprehensive Benchmarking Study
von: Liu, Che, et al.
Veröffentlicht: (2025)
von: Liu, Che, et al.
Veröffentlicht: (2025)
How Does Diverse Interpretability of Textual Prompts Impact Medical Vision-Language Zero-Shot Tasks?
von: Wang, Sicheng, et al.
Veröffentlicht: (2024)
von: Wang, Sicheng, et al.
Veröffentlicht: (2024)
Dynamic Decision Learning: Test-Time Evolution for Abnormality Grounding in Rare Diseases
von: Li, Jun, et al.
Veröffentlicht: (2026)
von: Li, Jun, et al.
Veröffentlicht: (2026)
Language Models Meet Anomaly Detection for Better Interpretability and Generalizability
von: Li, Jun, et al.
Veröffentlicht: (2024)
von: Li, Jun, et al.
Veröffentlicht: (2024)
IMITATE: Clinical Prior Guided Hierarchical Vision-Language Pre-training
von: Liu, Che, et al.
Veröffentlicht: (2023)
von: Liu, Che, et al.
Veröffentlicht: (2023)
G2D: From Global to Dense Radiography Representation Learning via Vision-Language Pre-training
von: Liu, Che, et al.
Veröffentlicht: (2023)
von: Liu, Che, et al.
Veröffentlicht: (2023)
BIMCV-R: A Landmark Dataset for 3D CT Text-Image Retrieval
von: Chen, Yinda, et al.
Veröffentlicht: (2024)
von: Chen, Yinda, et al.
Veröffentlicht: (2024)
Semantic Alignment of Unimodal Medical Text and Vision Representations
von: Di Folco, Maxime, et al.
Veröffentlicht: (2025)
von: Di Folco, Maxime, et al.
Veröffentlicht: (2025)
T3D: Advancing 3D Medical Vision-Language Pre-training by Learning Multi-View Visual Consistency
von: Liu, Che, et al.
Veröffentlicht: (2023)
von: Liu, Che, et al.
Veröffentlicht: (2023)
Denoising Diffusion Models for Anomaly Localization in Medical Images
von: Bercea, Cosmin I., et al.
Veröffentlicht: (2024)
von: Bercea, Cosmin I., et al.
Veröffentlicht: (2024)
Med-UniC: Unifying Cross-Lingual Medical Vision-Language Pre-Training by Diminishing Bias
von: Wan, Zhongwei, et al.
Veröffentlicht: (2023)
von: Wan, Zhongwei, et al.
Veröffentlicht: (2023)
Diffusion Models with Implicit Guidance for Medical Anomaly Detection
von: Bercea, Cosmin I., et al.
Veröffentlicht: (2024)
von: Bercea, Cosmin I., et al.
Veröffentlicht: (2024)
Interpretable Representation Learning of Cardiac MRI via Attribute Regularization
von: Di Folco, Maxime, et al.
Veröffentlicht: (2024)
von: Di Folco, Maxime, et al.
Veröffentlicht: (2024)
Argus: Benchmarking and Enhancing Vision-Language Models for 3D Radiology Report Generation
von: Liu, Che, et al.
Veröffentlicht: (2024)
von: Liu, Che, et al.
Veröffentlicht: (2024)
Towards Universal Unsupervised Anomaly Detection in Medical Imaging
von: Bercea, Cosmin I., et al.
Veröffentlicht: (2024)
von: Bercea, Cosmin I., et al.
Veröffentlicht: (2024)
LocBAM: Advancing 3D Patch-Based Image Segmentation by Integrating Location Contex
von: Hooft, Donnate, et al.
Veröffentlicht: (2026)
von: Hooft, Donnate, et al.
Veröffentlicht: (2026)
Freeze the backbones: A Parameter-Efficient Contrastive Approach to Robust Medical Vision-Language Pre-training
von: Qin, Jiuming, et al.
Veröffentlicht: (2024)
von: Qin, Jiuming, et al.
Veröffentlicht: (2024)
Does DINOv3 Set a New Medical Vision Standard? Benchmarking 2D and 3D Classification, Segmentation, and Registration
von: Liu, Che, et al.
Veröffentlicht: (2025)
von: Liu, Che, et al.
Veröffentlicht: (2025)
Influence of Classification Task and Distribution Shift Type on OOD Detection in Fetal Ultrasound
von: Wong, Chun Kit, et al.
Veröffentlicht: (2025)
von: Wong, Chun Kit, et al.
Veröffentlicht: (2025)
GroundingGPT:Language Enhanced Multi-modal Grounding Model
von: Li, Zhaowei, et al.
Veröffentlicht: (2024)
von: Li, Zhaowei, et al.
Veröffentlicht: (2024)
MedEdit: Counterfactual Diffusion-based Image Editing on Brain MRI
von: Alaya, Malek Ben, et al.
Veröffentlicht: (2024)
von: Alaya, Malek Ben, et al.
Veröffentlicht: (2024)
Light Up the Shadows: Enhance Long-Tailed Entity Grounding with Concept-Guided Vision-Language Models
von: Zhang, Yikai, et al.
Veröffentlicht: (2024)
von: Zhang, Yikai, et al.
Veröffentlicht: (2024)
Show and Guide: Instructional-Plan Grounded Vision and Language Model
von: Glória-Silva, Diogo, et al.
Veröffentlicht: (2024)
von: Glória-Silva, Diogo, et al.
Veröffentlicht: (2024)
Evaluation and Enhancement of Semantic Grounding in Large Vision-Language Models
von: Lu, Jiaying, et al.
Veröffentlicht: (2023)
von: Lu, Jiaying, et al.
Veröffentlicht: (2023)
Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models
von: Zhu, Tinghui, et al.
Veröffentlicht: (2024)
von: Zhu, Tinghui, et al.
Veröffentlicht: (2024)
Vision-Language Modeling in PET/CT for Visual Grounding of Positive Findings
von: Huemann, Zachary, et al.
Veröffentlicht: (2025)
von: Huemann, Zachary, et al.
Veröffentlicht: (2025)
VividMed: Vision Language Model with Versatile Visual Grounding for Medicine
von: Luo, Lingxiao, et al.
Veröffentlicht: (2024)
von: Luo, Lingxiao, et al.
Veröffentlicht: (2024)
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos
von: Hannan, Tanveer, et al.
Veröffentlicht: (2024)
von: Hannan, Tanveer, et al.
Veröffentlicht: (2024)
Beyond the Vision Encoder: Identifying and Mitigating Spatial Bias in Large Vision-Language Models
von: Zhu, Yingjie, et al.
Veröffentlicht: (2025)
von: Zhu, Yingjie, et al.
Veröffentlicht: (2025)
Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning
von: Zhu, Yingjie, et al.
Veröffentlicht: (2024)
von: Zhu, Yingjie, et al.
Veröffentlicht: (2024)
Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models
von: Gröpl, Marcel, et al.
Veröffentlicht: (2026)
von: Gröpl, Marcel, et al.
Veröffentlicht: (2026)
A Joint Study of Phrase Grounding and Task Performance in Vision and Language Models
von: Kojima, Noriyuki, et al.
Veröffentlicht: (2023)
von: Kojima, Noriyuki, et al.
Veröffentlicht: (2023)
MAGIC: Multimodal Alignment & Grounding-aware Instruction Coreset for Vision-Language Models
von: Biswas, Shristi Das, et al.
Veröffentlicht: (2026)
von: Biswas, Shristi Das, et al.
Veröffentlicht: (2026)
Does Object Grounding Really Reduce Hallucination of Large Vision-Language Models?
von: Geigle, Gregor, et al.
Veröffentlicht: (2024)
von: Geigle, Gregor, et al.
Veröffentlicht: (2024)
Expanding the Boundaries of Vision Prior Knowledge in Multi-modal Large Language Models
von: Liang, Qiao, et al.
Veröffentlicht: (2025)
von: Liang, Qiao, et al.
Veröffentlicht: (2025)
Can Medical Vision-Language Pre-training Succeed with Purely Synthetic Data?
von: Liu, Che, et al.
Veröffentlicht: (2024)
von: Liu, Che, et al.
Veröffentlicht: (2024)
"See the World, Discover Knowledge": A Chinese Factuality Evaluation for Large Vision Language Models
von: Gu, Jihao, et al.
Veröffentlicht: (2025)
von: Gu, Jihao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Knowledge to Sight: Reasoning over Visual Attributes via Knowledge Decomposition for Abnormality Grounding
von: Li, Jun, et al.
Veröffentlicht: (2025) -
FMBench: Benchmarking Fairness in Multimodal Large Language Models on Medical Tasks
von: Wu, Peiran, et al.
Veröffentlicht: (2024) -
Utilizing Synthetic Data for Medical Vision-Language Pre-training: Bypassing the Need for Real Images
von: Liu, Che, et al.
Veröffentlicht: (2023) -
How Far Have Medical Vision-Language Models Come? A Comprehensive Benchmarking Study
von: Liu, Che, et al.
Veröffentlicht: (2025) -
How Does Diverse Interpretability of Textual Prompts Impact Medical Vision-Language Zero-Shot Tasks?
von: Wang, Sicheng, et al.
Veröffentlicht: (2024)