Data Alignment for Zero-Shot Concept Generation in Dermatology AI
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gadgil, Soham, Bigverdi, Mahtab |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
No "Zero-Shot" Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model Performance
von: Udandarao, Vishaal, et al.
Veröffentlicht: (2024)
von: Udandarao, Vishaal, et al.
Veröffentlicht: (2024)
Image-Caption Encoding for Improving Zero-Shot Generalization
von: Yu, Eric Yang, et al.
Veröffentlicht: (2024)
von: Yu, Eric Yang, et al.
Veröffentlicht: (2024)
RadZero: Similarity-Based Cross-Attention for Explainable Vision-Language Alignment in Chest X-ray with Zero-Shot Multi-Task Capability
von: Park, Jonggwon, et al.
Veröffentlicht: (2025)
von: Park, Jonggwon, et al.
Veröffentlicht: (2025)
Classification for everyone : Building geography agnostic models for fairer recognition
von: Jindal, Akshat, et al.
Veröffentlicht: (2023)
von: Jindal, Akshat, et al.
Veröffentlicht: (2023)
Zero-Shot Defense Against Toxic Images via Inherent Multimodal Alignment in LVLMs
von: Zhao, Wei, et al.
Veröffentlicht: (2025)
von: Zhao, Wei, et al.
Veröffentlicht: (2025)
SurrogateSHAP: Training-Free Contributor Attribution for Text-to-Image (T2I) Models
von: Lu, Mingyu, et al.
Veröffentlicht: (2026)
von: Lu, Mingyu, et al.
Veröffentlicht: (2026)
Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering
von: Beliaev, Mark, et al.
Veröffentlicht: (2025)
von: Beliaev, Mark, et al.
Veröffentlicht: (2025)
D3G: Diverse Demographic Data Generation Increases Zero-Shot Image Classification Accuracy within Multimodal Models
von: Hickmon, Javon
Veröffentlicht: (2025)
von: Hickmon, Javon
Veröffentlicht: (2025)
FocalLens: Instruction Tuning Enables Zero-Shot Conditional Image Representations
von: Hsieh, Cheng-Yu, et al.
Veröffentlicht: (2025)
von: Hsieh, Cheng-Yu, et al.
Veröffentlicht: (2025)
Troika: Multi-Path Cross-Modal Traction for Compositional Zero-Shot Learning
von: Huang, Siteng, et al.
Veröffentlicht: (2023)
von: Huang, Siteng, et al.
Veröffentlicht: (2023)
CLIP meets DINO for Tuning Zero-Shot Classifier using Unlabeled Image Collections
von: Imam, Mohamed Fazli, et al.
Veröffentlicht: (2024)
von: Imam, Mohamed Fazli, et al.
Veröffentlicht: (2024)
Exploring the Limits of Zero Shot Vision Language Models for Hate Meme Detection: The Vulnerabilities and their Interpretations
von: Rizwan, Naquee, et al.
Veröffentlicht: (2024)
von: Rizwan, Naquee, et al.
Veröffentlicht: (2024)
CPLIP: Zero-Shot Learning for Histopathology with Comprehensive Vision-Language Alignment
von: Javed, Sajid, et al.
Veröffentlicht: (2024)
von: Javed, Sajid, et al.
Veröffentlicht: (2024)
Ablate-to-Validate: Are Vision-Language Models Really Using Continuous Thought Tokens?
von: Zhang, Tianyi, et al.
Veröffentlicht: (2026)
von: Zhang, Tianyi, et al.
Veröffentlicht: (2026)
Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning
von: Luo, Jianjie, et al.
Veröffentlicht: (2024)
von: Luo, Jianjie, et al.
Veröffentlicht: (2024)
Zero-Shot Vehicle Model Recognition via Text-Based Retrieval-Augmented Generation
von: Chang, Wei-Chia, et al.
Veröffentlicht: (2025)
von: Chang, Wei-Chia, et al.
Veröffentlicht: (2025)
FilterRAG: Zero-Shot Informed Retrieval-Augmented Generation to Mitigate Hallucinations in VQA
von: Sarwar, Nobin
Veröffentlicht: (2025)
von: Sarwar, Nobin
Veröffentlicht: (2025)
Verbalized Representation Learning for Interpretable Few-Shot Generalization
von: Yang, Cheng-Fu, et al.
Veröffentlicht: (2024)
von: Yang, Cheng-Fu, et al.
Veröffentlicht: (2024)
A Survey on Generative Modeling with Limited Data, Few Shots, and Zero Shot
von: Abdollahzadeh, Milad, et al.
Veröffentlicht: (2023)
von: Abdollahzadeh, Milad, et al.
Veröffentlicht: (2023)
Phrase-Instance Alignment for Generalized Referring Segmentation
von: Nguyen, E-Ro, et al.
Veröffentlicht: (2024)
von: Nguyen, E-Ro, et al.
Veröffentlicht: (2024)
Commonsense for Zero-Shot Natural Language Video Localization
von: Holla, Meghana, et al.
Veröffentlicht: (2023)
von: Holla, Meghana, et al.
Veröffentlicht: (2023)
Unified Vision-Language Modeling via Concept Space Alignment
von: Qiu, Yifu, et al.
Veröffentlicht: (2026)
von: Qiu, Yifu, et al.
Veröffentlicht: (2026)
Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputs
von: Shukor, Mustafa, et al.
Veröffentlicht: (2024)
von: Shukor, Mustafa, et al.
Veröffentlicht: (2024)
Zero-Shot Refinement of Buildings' Segmentation Models using SAM
von: Mayladan, Ali, et al.
Veröffentlicht: (2023)
von: Mayladan, Ali, et al.
Veröffentlicht: (2023)
ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion Models
von: Patel, Maitreya, et al.
Veröffentlicht: (2023)
von: Patel, Maitreya, et al.
Veröffentlicht: (2023)
Zero-Shot Visual Reasoning by Vision-Language Models: Benchmarking and Analysis
von: Nagar, Aishik, et al.
Veröffentlicht: (2024)
von: Nagar, Aishik, et al.
Veröffentlicht: (2024)
Distributionally Robust Alignment for Medical Federated Vision-Language Pre-training Under Data Heterogeneity
von: Shuai, Zitao, et al.
Veröffentlicht: (2024)
von: Shuai, Zitao, et al.
Veröffentlicht: (2024)
Benchmarking Zero-Shot Robustness of Multimodal Foundation Models: A Pilot Study
von: Wang, Chenguang, et al.
Veröffentlicht: (2024)
von: Wang, Chenguang, et al.
Veröffentlicht: (2024)
Perception Tokens Enhance Visual Reasoning in Multimodal Language Models
von: Bigverdi, Mahtab, et al.
Veröffentlicht: (2024)
von: Bigverdi, Mahtab, et al.
Veröffentlicht: (2024)
Do Vision and Language Models Share Concepts? A Vector Space Alignment Study
von: Li, Jiaang, et al.
Veröffentlicht: (2023)
von: Li, Jiaang, et al.
Veröffentlicht: (2023)
Multimodal Arabic Captioning with Interpretable Visual Concept Integration
von: Elchafei, Passant, et al.
Veröffentlicht: (2025)
von: Elchafei, Passant, et al.
Veröffentlicht: (2025)
Zero-Shot Dynamic Concept Personalization with Grid-Based LoRA
von: Abdal, Rameen, et al.
Veröffentlicht: (2025)
von: Abdal, Rameen, et al.
Veröffentlicht: (2025)
Controlled Training Data Generation with Diffusion Models
von: Yeo, Teresa, et al.
Veröffentlicht: (2024)
von: Yeo, Teresa, et al.
Veröffentlicht: (2024)
Data Redaction from Conditional Generative Models
von: Kong, Zhifeng, et al.
Veröffentlicht: (2023)
von: Kong, Zhifeng, et al.
Veröffentlicht: (2023)
Energy-based Models are Zero-Shot Planners for Compositional Scene Rearrangement
von: Gkanatsios, Nikolaos, et al.
Veröffentlicht: (2023)
von: Gkanatsios, Nikolaos, et al.
Veröffentlicht: (2023)
Text-centric Alignment for Multi-Modality Learning
von: Tsai, Yun-Da, et al.
Veröffentlicht: (2024)
von: Tsai, Yun-Da, et al.
Veröffentlicht: (2024)
VidLA: Video-Language Alignment at Scale
von: Rizve, Mamshad Nayeem, et al.
Veröffentlicht: (2024)
von: Rizve, Mamshad Nayeem, et al.
Veröffentlicht: (2024)
Listen Then See: Video Alignment with Speaker Attention
von: Agrawal, Aviral, et al.
Veröffentlicht: (2024)
von: Agrawal, Aviral, et al.
Veröffentlicht: (2024)
Exploring Multi-Grained Concept Annotations for Multimodal Large Language Models
von: Xu, Xiao, et al.
Veröffentlicht: (2024)
von: Xu, Xiao, et al.
Veröffentlicht: (2024)
When are Lemons Purple? The Concept Association Bias of Vision-Language Models
von: Yamada, Yutaro, et al.
Veröffentlicht: (2022)
von: Yamada, Yutaro, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
No "Zero-Shot" Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model Performance
von: Udandarao, Vishaal, et al.
Veröffentlicht: (2024) -
Image-Caption Encoding for Improving Zero-Shot Generalization
von: Yu, Eric Yang, et al.
Veröffentlicht: (2024) -
RadZero: Similarity-Based Cross-Attention for Explainable Vision-Language Alignment in Chest X-ray with Zero-Shot Multi-Task Capability
von: Park, Jonggwon, et al.
Veröffentlicht: (2025) -
Classification for everyone : Building geography agnostic models for fairer recognition
von: Jindal, Akshat, et al.
Veröffentlicht: (2023) -
Zero-Shot Defense Against Toxic Images via Inherent Multimodal Alignment in LVLMs
von: Zhao, Wei, et al.
Veröffentlicht: (2025)