LLaVA Needs More Knowledge: Retrieval Augmented Natural Language Generation with Knowledge Graph for Explaining Thoracic Pathologies
Fuente:
arXiv
Salvato in:
| Autori principali: | Hamza, Ameer, Abdullah, Ahn, Yong Hyun, Lee, Sungyoung, Kim, Seong Tae |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Resource-Efficient Medical Report Generation using Large Language Models
di: Abdullah, et al.
Pubblicazione: (2024)
di: Abdullah, et al.
Pubblicazione: (2024)
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
di: Shu, Fangxun, et al.
Pubblicazione: (2024)
di: Shu, Fangxun, et al.
Pubblicazione: (2024)
VLM-KG: Multimodal Radiology Knowledge Graph Generation
di: Abdullah, Abdullah, et al.
Pubblicazione: (2025)
di: Abdullah, Abdullah, et al.
Pubblicazione: (2025)
Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs
di: Caffagni, Davide, et al.
Pubblicazione: (2024)
di: Caffagni, Davide, et al.
Pubblicazione: (2024)
ViDRiP-LLaVA: A Dataset and Benchmark for Diagnostic Reasoning from Pathology Videos
di: Vuong, Trinh T. L., et al.
Pubblicazione: (2025)
di: Vuong, Trinh T. L., et al.
Pubblicazione: (2025)
LLaVA-CKD: Bottom-Up Cascaded Knowledge Distillation for Vision-Language Models
di: Gkalelis, Nikolaos, et al.
Pubblicazione: (2026)
di: Gkalelis, Nikolaos, et al.
Pubblicazione: (2026)
Cosmos-LLaVA: Chatting with the Visual Cosmos-LLaVA: Görselle Sohbet Etmek
di: Zeer, Ahmed, et al.
Pubblicazione: (2024)
di: Zeer, Ahmed, et al.
Pubblicazione: (2024)
TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings
di: Yan, Dawei, et al.
Pubblicazione: (2024)
di: Yan, Dawei, et al.
Pubblicazione: (2024)
WWW: A Unified Framework for Explaining What, Where and Why of Neural Networks by Interpretation of Neuron Concepts
di: Ahn, Yong Hyun, et al.
Pubblicazione: (2024)
di: Ahn, Yong Hyun, et al.
Pubblicazione: (2024)
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval
di: Lu, Weiheng, et al.
Pubblicazione: (2024)
di: Lu, Weiheng, et al.
Pubblicazione: (2024)
PA-LLaVA: A Large Language-Vision Assistant for Human Pathology Image Understanding
di: Dai, Dawei, et al.
Pubblicazione: (2024)
di: Dai, Dawei, et al.
Pubblicazione: (2024)
Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases
di: Wang, Liqiong, et al.
Pubblicazione: (2024)
di: Wang, Liqiong, et al.
Pubblicazione: (2024)
X-LLaVA: Optimizing Bilingual Large Vision-Language Alignment
di: Shin, Dongjae, et al.
Pubblicazione: (2024)
di: Shin, Dongjae, et al.
Pubblicazione: (2024)
Amharic LLaMA and LLaVA: Multimodal LLMs for Low Resource Languages
di: Andersland, Michael
Pubblicazione: (2024)
di: Andersland, Michael
Pubblicazione: (2024)
Normal and Abnormal Pathology Knowledge-Augmented Vision-Language Model for Anomaly Detection in Pathology Images
di: Song, Jinsol, et al.
Pubblicazione: (2025)
di: Song, Jinsol, et al.
Pubblicazione: (2025)
LLaVA-SLT: Visual Language Tuning for Sign Language Translation
di: Liang, Han, et al.
Pubblicazione: (2024)
di: Liang, Han, et al.
Pubblicazione: (2024)
LLaVA-VSD: Large Language-and-Vision Assistant for Visual Spatial Description
di: Jin, Yizhang, et al.
Pubblicazione: (2024)
di: Jin, Yizhang, et al.
Pubblicazione: (2024)
LLaVA-KD: A Framework of Distilling Multimodal Large Language Models
di: Cai, Yuxuan, et al.
Pubblicazione: (2024)
di: Cai, Yuxuan, et al.
Pubblicazione: (2024)
Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models
di: Shi, Wenhao, et al.
Pubblicazione: (2024)
di: Shi, Wenhao, et al.
Pubblicazione: (2024)
LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models
di: Shang, Yuzhang, et al.
Pubblicazione: (2024)
di: Shang, Yuzhang, et al.
Pubblicazione: (2024)
LLaVA-Docent: Instruction Tuning with Multimodal Large Language Model to Support Art Appreciation Education
di: Lee, Unggi, et al.
Pubblicazione: (2024)
di: Lee, Unggi, et al.
Pubblicazione: (2024)
SQ-LLaVA: Self-Questioning for Large Vision-Language Assistant
di: Sun, Guohao, et al.
Pubblicazione: (2024)
di: Sun, Guohao, et al.
Pubblicazione: (2024)
LLaVA-Ultra: Large Chinese Language and Vision Assistant for Ultrasound
di: Guo, Xuechen, et al.
Pubblicazione: (2024)
di: Guo, Xuechen, et al.
Pubblicazione: (2024)
LLaVA-Read: Enhancing Reading Ability of Multimodal Language Models
di: Zhang, Ruiyi, et al.
Pubblicazione: (2024)
di: Zhang, Ruiyi, et al.
Pubblicazione: (2024)
MC-LLaVA: Multi-Concept Personalized Vision-Language Model
di: An, Ruichuan, et al.
Pubblicazione: (2024)
di: An, Ruichuan, et al.
Pubblicazione: (2024)
MC-LLaVA: Multi-Concept Personalized Vision-Language Model
di: An, Ruichuan, et al.
Pubblicazione: (2025)
di: An, Ruichuan, et al.
Pubblicazione: (2025)
LLaVA-LE: Large Language-and-Vision Assistant for Lunar Exploration
di: Inal, Gokce, et al.
Pubblicazione: (2026)
di: Inal, Gokce, et al.
Pubblicazione: (2026)
LLaVA-SG: Leveraging Scene Graphs as Visual Semantic Expression in Vision-Language Models
di: Wang, Jingyi, et al.
Pubblicazione: (2024)
di: Wang, Jingyi, et al.
Pubblicazione: (2024)
LLaVA-Critic: Learning to Evaluate Multimodal Models
di: Xiong, Tianyi, et al.
Pubblicazione: (2024)
di: Xiong, Tianyi, et al.
Pubblicazione: (2024)
LLaVAC: Fine-tuning LLaVA as a Multimodal Sentiment Classifier
di: Chay-intr, T., et al.
Pubblicazione: (2025)
di: Chay-intr, T., et al.
Pubblicazione: (2025)
Retrieval-Augmented Natural Language Reasoning for Explainable Visual Question Answering
di: Lim, Su Hyeon, et al.
Pubblicazione: (2024)
di: Lim, Su Hyeon, et al.
Pubblicazione: (2024)
Safe-LLaVA: A Privacy-Preserving Vision-Language Dataset and Benchmark for Biometric Safety
di: Kim, Younggun, et al.
Pubblicazione: (2025)
di: Kim, Younggun, et al.
Pubblicazione: (2025)
When LLaVA Meets Objects: Token Composition for Vision-Language-Models
di: Jahagirdar, Soumya, et al.
Pubblicazione: (2026)
di: Jahagirdar, Soumya, et al.
Pubblicazione: (2026)
LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model
di: Zhu, Yichen, et al.
Pubblicazione: (2024)
di: Zhu, Yichen, et al.
Pubblicazione: (2024)
Space-LLaVA: a Vision-Language Model Adapted to Extraterrestrial Applications
di: Foutter, Matthew, et al.
Pubblicazione: (2024)
di: Foutter, Matthew, et al.
Pubblicazione: (2024)
ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models
di: Ye, Xubing, et al.
Pubblicazione: (2024)
di: Ye, Xubing, et al.
Pubblicazione: (2024)
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
di: Lin, Bin, et al.
Pubblicazione: (2024)
di: Lin, Bin, et al.
Pubblicazione: (2024)
Why do LLaVA Vision-Language Models Reply to Images in English?
di: Hinck, Musashi, et al.
Pubblicazione: (2024)
di: Hinck, Musashi, et al.
Pubblicazione: (2024)
LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
di: Xu, Guowei, et al.
Pubblicazione: (2024)
di: Xu, Guowei, et al.
Pubblicazione: (2024)
Continual LLaVA: Continual Instruction Tuning in Large Vision-Language Models
di: Cao, Meng, et al.
Pubblicazione: (2024)
di: Cao, Meng, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Resource-Efficient Medical Report Generation using Large Language Models
di: Abdullah, et al.
Pubblicazione: (2024) -
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
di: Shu, Fangxun, et al.
Pubblicazione: (2024) -
VLM-KG: Multimodal Radiology Knowledge Graph Generation
di: Abdullah, Abdullah, et al.
Pubblicazione: (2025) -
Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs
di: Caffagni, Davide, et al.
Pubblicazione: (2024) -
ViDRiP-LLaVA: A Dataset and Benchmark for Diagnostic Reasoning from Pathology Videos
di: Vuong, Trinh T. L., et al.
Pubblicazione: (2025)