ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cai, Mu, Liu, Haotian, Park, Dennis, Mustikovela, Siva Karthik, Meyer, Gregory P., Chai, Yuning, Lee, Yong Jae |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models
von: Shang, Yuzhang, et al.
Veröffentlicht: (2024)
von: Shang, Yuzhang, et al.
Veröffentlicht: (2024)
VLMine: Long-Tail Data Mining with Vision Language Models
von: Ye, Mao, et al.
Veröffentlicht: (2024)
von: Ye, Mao, et al.
Veröffentlicht: (2024)
Cosmos-LLaVA: Chatting with the Visual Cosmos-LLaVA: Görselle Sohbet Etmek
von: Zeer, Ahmed, et al.
Veröffentlicht: (2024)
von: Zeer, Ahmed, et al.
Veröffentlicht: (2024)
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
von: Shu, Fangxun, et al.
Veröffentlicht: (2024)
von: Shu, Fangxun, et al.
Veröffentlicht: (2024)
LLaVA-Critic: Learning to Evaluate Multimodal Models
von: Xiong, Tianyi, et al.
Veröffentlicht: (2024)
von: Xiong, Tianyi, et al.
Veröffentlicht: (2024)
AVG-LLaVA: An Efficient Large Multimodal Model with Adaptive Visual Granularity
von: Lan, Zhibin, et al.
Veröffentlicht: (2024)
von: Lan, Zhibin, et al.
Veröffentlicht: (2024)
Dr-LLaVA: Visual Instruction Tuning with Symbolic Clinical Grounding
von: Sun, Shenghuan, et al.
Veröffentlicht: (2024)
von: Sun, Shenghuan, et al.
Veröffentlicht: (2024)
NOVO: Bridging LLaVA and SAM with Visual-only Prompts for Reasoning Segmentation
von: Yoon, Kyung-Yoon, et al.
Veröffentlicht: (2025)
von: Yoon, Kyung-Yoon, et al.
Veröffentlicht: (2025)
LLaVAC: Fine-tuning LLaVA as a Multimodal Sentiment Classifier
von: Chay-intr, T., et al.
Veröffentlicht: (2025)
von: Chay-intr, T., et al.
Veröffentlicht: (2025)
Amharic LLaMA and LLaVA: Multimodal LLMs for Low Resource Languages
von: Andersland, Michael
Veröffentlicht: (2024)
von: Andersland, Michael
Veröffentlicht: (2024)
TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings
von: Yan, Dawei, et al.
Veröffentlicht: (2024)
von: Yan, Dawei, et al.
Veröffentlicht: (2024)
Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA
von: Kanjula, Karthik Reddy, et al.
Veröffentlicht: (2025)
von: Kanjula, Karthik Reddy, et al.
Veröffentlicht: (2025)
Enhance Image-to-Image Generation with LLaVA-generated Prompts
von: Ding, Zhicheng, et al.
Veröffentlicht: (2024)
von: Ding, Zhicheng, et al.
Veröffentlicht: (2024)
LLaVA-OneVision: Easy Visual Task Transfer
von: Li, Bo, et al.
Veröffentlicht: (2024)
von: Li, Bo, et al.
Veröffentlicht: (2024)
LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs?
von: Fang, Kechen, et al.
Veröffentlicht: (2026)
von: Fang, Kechen, et al.
Veröffentlicht: (2026)
Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs
von: Caffagni, Davide, et al.
Veröffentlicht: (2024)
von: Caffagni, Davide, et al.
Veröffentlicht: (2024)
LLaVA-c: Continual Improved Visual Instruction Tuning
von: Liu, Wenzhuo, et al.
Veröffentlicht: (2025)
von: Liu, Wenzhuo, et al.
Veröffentlicht: (2025)
LLaVA-KD: A Framework of Distilling Multimodal Large Language Models
von: Cai, Yuxuan, et al.
Veröffentlicht: (2024)
von: Cai, Yuxuan, et al.
Veröffentlicht: (2024)
R-LLaVA: Improving Med-VQA Understanding through Visual Region of Interest
von: Chen, Xupeng, et al.
Veröffentlicht: (2024)
von: Chen, Xupeng, et al.
Veröffentlicht: (2024)
LLaVA-Ultra: Large Chinese Language and Vision Assistant for Ultrasound
von: Guo, Xuechen, et al.
Veröffentlicht: (2024)
von: Guo, Xuechen, et al.
Veröffentlicht: (2024)
ViDRiP-LLaVA: A Dataset and Benchmark for Diagnostic Reasoning from Pathology Videos
von: Vuong, Trinh T. L., et al.
Veröffentlicht: (2025)
von: Vuong, Trinh T. L., et al.
Veröffentlicht: (2025)
LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding
von: Zhou, Hanyu, et al.
Veröffentlicht: (2025)
von: Zhou, Hanyu, et al.
Veröffentlicht: (2025)
Face-LLaVA: Facial Expression and Attribute Understanding through Instruction Tuning
von: Chaubey, Ashutosh, et al.
Veröffentlicht: (2025)
von: Chaubey, Ashutosh, et al.
Veröffentlicht: (2025)
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs
von: Lou, Haoran, et al.
Veröffentlicht: (2025)
von: Lou, Haoran, et al.
Veröffentlicht: (2025)
Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models
von: Shi, Wenhao, et al.
Veröffentlicht: (2024)
von: Shi, Wenhao, et al.
Veröffentlicht: (2024)
LLaVA-SLT: Visual Language Tuning for Sign Language Translation
von: Liang, Han, et al.
Veröffentlicht: (2024)
von: Liang, Han, et al.
Veröffentlicht: (2024)
MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2024)
Yo'LLaVA: Your Personalized Language and Vision Assistant
von: Nguyen, Thao, et al.
Veröffentlicht: (2024)
von: Nguyen, Thao, et al.
Veröffentlicht: (2024)
WSI-LLaVA: A Multimodal Large Language Model for Whole Slide Image
von: Liang, Yuci, et al.
Veröffentlicht: (2024)
von: Liang, Yuci, et al.
Veröffentlicht: (2024)
LLaVA-NeuMT: Selective Layer-Neuron Modulation for Efficient Multilingual Multimodal Translation
von: Wei, Jingxuan, et al.
Veröffentlicht: (2025)
von: Wei, Jingxuan, et al.
Veröffentlicht: (2025)
LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding
von: Li, Hongyu, et al.
Veröffentlicht: (2025)
von: Li, Hongyu, et al.
Veröffentlicht: (2025)
LLaVA-Read: Enhancing Reading Ability of Multimodal Language Models
von: Zhang, Ruiyi, et al.
Veröffentlicht: (2024)
von: Zhang, Ruiyi, et al.
Veröffentlicht: (2024)
Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
von: Zhang, Yi-Fan, et al.
Veröffentlicht: (2024)
von: Zhang, Yi-Fan, et al.
Veröffentlicht: (2024)
LLaVA-FA: Learning Fourier Approximation for Compressing Large Multimodal Models
von: Zheng, Pengcheng, et al.
Veröffentlicht: (2026)
von: Zheng, Pengcheng, et al.
Veröffentlicht: (2026)
LLaVA-Video: Video Instruction Tuning With Synthetic Data
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2024)
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
von: Lin, Bin, et al.
Veröffentlicht: (2023)
von: Lin, Bin, et al.
Veröffentlicht: (2023)
Purrfessor: A Fine-tuned Multimodal LLaVA Diet Health Chatbot
von: Lu, Linqi, et al.
Veröffentlicht: (2024)
von: Lu, Linqi, et al.
Veröffentlicht: (2024)
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
von: Yuan, Haobo, et al.
Veröffentlicht: (2025)
von: Yuan, Haobo, et al.
Veröffentlicht: (2025)
LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
von: Zhang, Shaolei, et al.
Veröffentlicht: (2025)
von: Zhang, Shaolei, et al.
Veröffentlicht: (2025)
LLaVA-VSD: Large Language-and-Vision Assistant for Visual Spatial Description
von: Jin, Yizhang, et al.
Veröffentlicht: (2024)
von: Jin, Yizhang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models
von: Shang, Yuzhang, et al.
Veröffentlicht: (2024) -
VLMine: Long-Tail Data Mining with Vision Language Models
von: Ye, Mao, et al.
Veröffentlicht: (2024) -
Cosmos-LLaVA: Chatting with the Visual Cosmos-LLaVA: Görselle Sohbet Etmek
von: Zeer, Ahmed, et al.
Veröffentlicht: (2024) -
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
von: Shu, Fangxun, et al.
Veröffentlicht: (2024) -
LLaVA-Critic: Learning to Evaluate Multimodal Models
von: Xiong, Tianyi, et al.
Veröffentlicht: (2024)