Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Seyfioglu, Mehmet Saygin, Ikezogwo, Wisdom O., Ghezloo, Fatemeh, Krishna, Ranjay, Shapiro, Linda |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MedicalNarratives: Connecting Medical Vision and Language with Localized Narratives
von: Ikezogwo, Wisdom O., et al.
Veröffentlicht: (2025)
von: Ikezogwo, Wisdom O., et al.
Veröffentlicht: (2025)
Quilt-1M: One Million Image-Text Pairs for Histopathology
von: Ikezogwo, Wisdom Oluchi, et al.
Veröffentlicht: (2023)
von: Ikezogwo, Wisdom Oluchi, et al.
Veröffentlicht: (2023)
PathFinder: A Multi-Modal Multi-Agent System for Medical Diagnostic Decision-Making Applied to Histopathology
von: Ghezloo, Fatemeh, et al.
Veröffentlicht: (2025)
von: Ghezloo, Fatemeh, et al.
Veröffentlicht: (2025)
When Rubrics Fail: Error Enumeration as Reward in Reference-Free RL Post-Training for Virtual Try-On
von: Ikezogwo, Wisdom, et al.
Veröffentlicht: (2026)
von: Ikezogwo, Wisdom, et al.
Veröffentlicht: (2026)
LLaVA-Video: Video Instruction Tuning With Synthetic Data
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2024)
LLaVA-c: Continual Improved Visual Instruction Tuning
von: Liu, Wenzhuo, et al.
Veröffentlicht: (2025)
von: Liu, Wenzhuo, et al.
Veröffentlicht: (2025)
Dr-LLaVA: Visual Instruction Tuning with Symbolic Clinical Grounding
von: Sun, Shenghuan, et al.
Veröffentlicht: (2024)
von: Sun, Shenghuan, et al.
Veröffentlicht: (2024)
MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2024)
Cosmos-LLaVA: Chatting with the Visual Cosmos-LLaVA: Görselle Sohbet Etmek
von: Zeer, Ahmed, et al.
Veröffentlicht: (2024)
von: Zeer, Ahmed, et al.
Veröffentlicht: (2024)
LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning
von: Cocchi, Federico, et al.
Veröffentlicht: (2025)
von: Cocchi, Federico, et al.
Veröffentlicht: (2025)
Face-LLaVA: Facial Expression and Attribute Understanding through Instruction Tuning
von: Chaubey, Ashutosh, et al.
Veröffentlicht: (2025)
von: Chaubey, Ashutosh, et al.
Veröffentlicht: (2025)
Continual LLaVA: Continual Instruction Tuning in Large Vision-Language Models
von: Cao, Meng, et al.
Veröffentlicht: (2024)
von: Cao, Meng, et al.
Veröffentlicht: (2024)
LLaVA-SLT: Visual Language Tuning for Sign Language Translation
von: Liang, Han, et al.
Veröffentlicht: (2024)
von: Liang, Han, et al.
Veröffentlicht: (2024)
LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering
von: Bi, Jinhe, et al.
Veröffentlicht: (2024)
von: Bi, Jinhe, et al.
Veröffentlicht: (2024)
TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings
von: Yan, Dawei, et al.
Veröffentlicht: (2024)
von: Yan, Dawei, et al.
Veröffentlicht: (2024)
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
von: Lin, Bin, et al.
Veröffentlicht: (2023)
von: Lin, Bin, et al.
Veröffentlicht: (2023)
LLaVA-OneVision: Easy Visual Task Transfer
von: Li, Bo, et al.
Veröffentlicht: (2024)
von: Li, Bo, et al.
Veröffentlicht: (2024)
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
von: Shu, Fangxun, et al.
Veröffentlicht: (2024)
von: Shu, Fangxun, et al.
Veröffentlicht: (2024)
LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding
von: Sun, Boyuan, et al.
Veröffentlicht: (2025)
von: Sun, Boyuan, et al.
Veröffentlicht: (2025)
LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations
von: Xu, Mingjie, et al.
Veröffentlicht: (2024)
von: Xu, Mingjie, et al.
Veröffentlicht: (2024)
HiDe-LLaVA: Hierarchical Decoupling for Continual Instruction Tuning of Multimodal Large Language Model
von: Guo, Haiyang, et al.
Veröffentlicht: (2025)
von: Guo, Haiyang, et al.
Veröffentlicht: (2025)
LLaVA-Pose: Enhancing Human Pose and Action Understanding via Keypoint-Integrated Instruction Tuning
von: Zhang, Dewen, et al.
Veröffentlicht: (2025)
von: Zhang, Dewen, et al.
Veröffentlicht: (2025)
Synthetic Visual Genome 2: Extracting Large-scale Spatio-Temporal Scene Graphs from Videos
von: Gao, Ziqi, et al.
Veröffentlicht: (2026)
von: Gao, Ziqi, et al.
Veröffentlicht: (2026)
LLaVA-Critic: Learning to Evaluate Multimodal Models
von: Xiong, Tianyi, et al.
Veröffentlicht: (2024)
von: Xiong, Tianyi, et al.
Veröffentlicht: (2024)
LLaVA-Docent: Instruction Tuning with Multimodal Large Language Model to Support Art Appreciation Education
von: Lee, Unggi, et al.
Veröffentlicht: (2024)
von: Lee, Unggi, et al.
Veröffentlicht: (2024)
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs
von: Lou, Haoran, et al.
Veröffentlicht: (2025)
von: Lou, Haoran, et al.
Veröffentlicht: (2025)
AVG-LLaVA: An Efficient Large Multimodal Model with Adaptive Visual Granularity
von: Lan, Zhibin, et al.
Veröffentlicht: (2024)
von: Lan, Zhibin, et al.
Veröffentlicht: (2024)
Spatial Reasoning is Not a Free Lunch: A Controlled Study on LLaVA
von: Alam, Nahid, et al.
Veröffentlicht: (2026)
von: Alam, Nahid, et al.
Veröffentlicht: (2026)
LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs
von: Sun, Boyuan, et al.
Veröffentlicht: (2025)
von: Sun, Boyuan, et al.
Veröffentlicht: (2025)
LLaVAC: Fine-tuning LLaVA as a Multimodal Sentiment Classifier
von: Chay-intr, T., et al.
Veröffentlicht: (2025)
von: Chay-intr, T., et al.
Veröffentlicht: (2025)
Amharic LLaMA and LLaVA: Multimodal LLMs for Low Resource Languages
von: Andersland, Michael
Veröffentlicht: (2024)
von: Andersland, Michael
Veröffentlicht: (2024)
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
von: Xu, Lin, et al.
Veröffentlicht: (2024)
von: Xu, Lin, et al.
Veröffentlicht: (2024)
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval
von: Lu, Weiheng, et al.
Veröffentlicht: (2024)
von: Lu, Weiheng, et al.
Veröffentlicht: (2024)
NOVO: Bridging LLaVA and SAM with Visual-only Prompts for Reasoning Segmentation
von: Yoon, Kyung-Yoon, et al.
Veröffentlicht: (2025)
von: Yoon, Kyung-Yoon, et al.
Veröffentlicht: (2025)
LLaVA-VSD: Large Language-and-Vision Assistant for Visual Spatial Description
von: Jin, Yizhang, et al.
Veröffentlicht: (2024)
von: Jin, Yizhang, et al.
Veröffentlicht: (2024)
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
von: Yuan, Haobo, et al.
Veröffentlicht: (2025)
von: Yuan, Haobo, et al.
Veröffentlicht: (2025)
LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
von: Zhang, Shaolei, et al.
Veröffentlicht: (2025)
von: Zhang, Shaolei, et al.
Veröffentlicht: (2025)
LLaVA-MLB: Mitigating and Leveraging Attention Bias for Training-Free Video LLMs
von: Shen, Leqi, et al.
Veröffentlicht: (2025)
von: Shen, Leqi, et al.
Veröffentlicht: (2025)
LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs?
von: Fang, Kechen, et al.
Veröffentlicht: (2026)
von: Fang, Kechen, et al.
Veröffentlicht: (2026)
ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts
von: Cai, Mu, et al.
Veröffentlicht: (2023)
von: Cai, Mu, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
MedicalNarratives: Connecting Medical Vision and Language with Localized Narratives
von: Ikezogwo, Wisdom O., et al.
Veröffentlicht: (2025) -
Quilt-1M: One Million Image-Text Pairs for Histopathology
von: Ikezogwo, Wisdom Oluchi, et al.
Veröffentlicht: (2023) -
PathFinder: A Multi-Modal Multi-Agent System for Medical Diagnostic Decision-Making Applied to Histopathology
von: Ghezloo, Fatemeh, et al.
Veröffentlicht: (2025) -
When Rubrics Fail: Error Enumeration as Reward in Reference-Free RL Post-Training for Virtual Try-On
von: Ikezogwo, Wisdom, et al.
Veröffentlicht: (2026) -
LLaVA-Video: Video Instruction Tuning With Synthetic Data
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2024)