Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Honglin, Gao, Yuting, Zhu, Chenglu, Chen, Jingdong, Yang, Ming, Yang, Lin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Accelerating Pre-training of Multimodal LLMs via Chain-of-Sight
by: Huang, Ziyuan, et al.
Published: (2024)
by: Huang, Ziyuan, et al.
Published: (2024)
WSI-VQA: Interpreting Whole Slide Images by Generative Visual Question Answering
by: Chen, Pingyi, et al.
Published: (2024)
by: Chen, Pingyi, et al.
Published: (2024)
PathVQ: Reforming Computational Pathology Foundation Model for Whole Slide Image Analysis via Vector Quantization
by: Li, Honglin, et al.
Published: (2025)
by: Li, Honglin, et al.
Published: (2025)
Rethinking Transformer for Long Contextual Histopathology Whole Slide Image Analysis
by: Li, Honglin, et al.
Published: (2024)
by: Li, Honglin, et al.
Published: (2024)
Attention-Challenging Multiple Instance Learning for Whole Slide Image Classification
by: Zhang, Yunlong, et al.
Published: (2023)
by: Zhang, Yunlong, et al.
Published: (2023)
Unleashing the Power of Prompt-driven Nucleus Instance Segmentation
by: Shui, Zhongyi, et al.
Published: (2023)
by: Shui, Zhongyi, et al.
Published: (2023)
Instruction Tuning-free Visual Token Complement for Multimodal LLMs
by: Wang, Dongsheng, et al.
Published: (2024)
by: Wang, Dongsheng, et al.
Published: (2024)
ModalPrompt: Towards Efficient Multimodal Continual Instruction Tuning with Dual-Modality Guided Prompt
by: Zeng, Fanhu, et al.
Published: (2024)
by: Zeng, Fanhu, et al.
Published: (2024)
MotionSight: Boosting Fine-Grained Motion Understanding in Multimodal LLMs
by: Du, Yipeng, et al.
Published: (2025)
by: Du, Yipeng, et al.
Published: (2025)
AEM: Attention Entropy Maximization for Multiple Instance Learning based Whole Slide Image Classification
by: Zhang, Yunlong, et al.
Published: (2024)
by: Zhang, Yunlong, et al.
Published: (2024)
WsiCaption: Multiple Instance Generation of Pathology Reports for Gigapixel Whole-Slide Images
by: Chen, Pingyi, et al.
Published: (2023)
by: Chen, Pingyi, et al.
Published: (2023)
Hints of Prompt: Enhancing Visual Representation for Multimodal LLMs in Autonomous Driving
by: Zhou, Hao, et al.
Published: (2024)
by: Zhou, Hao, et al.
Published: (2024)
Guiding Visual Autoregressive Models through Spectrum Weakening
by: Wang, Chaoyang, et al.
Published: (2025)
by: Wang, Chaoyang, et al.
Published: (2025)
Elevating Visual Perception in Multimodal LLMs with Visual Embedding Distillation
by: Jain, Jitesh, et al.
Published: (2024)
by: Jain, Jitesh, et al.
Published: (2024)
Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers
by: Yao, Yuxuan, et al.
Published: (2026)
by: Yao, Yuxuan, et al.
Published: (2026)
Mask-Guided Attention Regulation for Anatomically Consistent Counterfactual CXR Synthesis
by: Zhang, Zichun, et al.
Published: (2026)
by: Zhang, Zichun, et al.
Published: (2026)
PathAsst: A Generative Foundation AI Assistant Towards Artificial General Intelligence of Pathology
by: Sun, Yuxuan, et al.
Published: (2023)
by: Sun, Yuxuan, et al.
Published: (2023)
VRP-SAM: SAM with Visual Reference Prompt
by: Sun, Yanpeng, et al.
Published: (2024)
by: Sun, Yanpeng, et al.
Published: (2024)
IG-Diff: Complex Night Scene Restoration with Illumination-Guided Diffusion Model
by: Chen, Yifan, et al.
Published: (2026)
by: Chen, Yifan, et al.
Published: (2026)
VP-MEL: Visual Prompts Guided Multimodal Entity Linking
by: Mi, Hongze, et al.
Published: (2024)
by: Mi, Hongze, et al.
Published: (2024)
Large-scale cervical precancerous screening via AI-assisted cytology whole slide image analysis
by: Li, Honglin, et al.
Published: (2024)
by: Li, Honglin, et al.
Published: (2024)
Visual Instruction Pretraining for Domain-Specific Foundation Models
by: Li, Yuxuan, et al.
Published: (2025)
by: Li, Yuxuan, et al.
Published: (2025)
Visual Prompting in LLMs for Enhancing Emotion Recognition
by: Zhang, Qixuan, et al.
Published: (2024)
by: Zhang, Qixuan, et al.
Published: (2024)
CauSight: Learning to Supersense for Visual Causal Discovery
by: Zhang, Yize, et al.
Published: (2025)
by: Zhang, Yize, et al.
Published: (2025)
Insight Over Sight: Exploring the Vision-Knowledge Conflicts in Multimodal LLMs
by: Liu, Xiaoyuan, et al.
Published: (2024)
by: Liu, Xiaoyuan, et al.
Published: (2024)
Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception
by: Sun, Yanpeng, et al.
Published: (2024)
by: Sun, Yanpeng, et al.
Published: (2024)
Image-to-Brain Signal Generation for Visual Prosthesis with CLIP Guided Multimodal Diffusion Models
by: Xu, Ganxi, et al.
Published: (2025)
by: Xu, Ganxi, et al.
Published: (2025)
MGIMM: Multi-Granularity Instruction Multimodal Model for Attribute-Guided Remote Sensing Image Detailed Description
by: Yang, Cong, et al.
Published: (2024)
by: Yang, Cong, et al.
Published: (2024)
Guiding Instruction-based Image Editing via Multimodal Large Language Models
by: Fu, Tsu-Jui, et al.
Published: (2023)
by: Fu, Tsu-Jui, et al.
Published: (2023)
CPath-Omni: A Unified Multimodal Foundation Model for Patch and Whole Slide Image Analysis in Computational Pathology
by: Sun, Yuxuan, et al.
Published: (2024)
by: Sun, Yuxuan, et al.
Published: (2024)
Customized Visual Storytelling with Unified Multimodal LLMs
by: Li, Wei-Hua, et al.
Published: (2026)
by: Li, Wei-Hua, et al.
Published: (2026)
Pyramidal Patchification Flow for Visual Generation
by: Li, Hui, et al.
Published: (2025)
by: Li, Hui, et al.
Published: (2025)
INST-IT: Boosting Instance Understanding via Explicit Visual Prompt Instruction Tuning
by: Peng, Wujian, et al.
Published: (2024)
by: Peng, Wujian, et al.
Published: (2024)
Improving Generalized Visual Grounding with Instance-aware Joint Learning
by: Dai, Ming, et al.
Published: (2025)
by: Dai, Ming, et al.
Published: (2025)
OwlSight: A Robust Illumination Adaptation Framework for Dark Video Human Action Recognition
by: Cheng, Shihao, et al.
Published: (2025)
by: Cheng, Shihao, et al.
Published: (2025)
UniAPO: Unified Multimodal Automated Prompt Optimization
by: Zhu, Qipeng, et al.
Published: (2025)
by: Zhu, Qipeng, et al.
Published: (2025)
EAGLE: Towards Efficient Arbitrary Referring Visual Prompts Comprehension for Multimodal Large Language Models
by: Zhang, Jiacheng, et al.
Published: (2024)
by: Zhang, Jiacheng, et al.
Published: (2024)
DA-VPT: Semantic-Guided Visual Prompt Tuning for Vision Transformers
by: Ren, Li, et al.
Published: (2025)
by: Ren, Li, et al.
Published: (2025)
EM-KD: Distilling Efficient Multimodal Large Language Model with Unbalanced Vision Tokens
by: Feng, Ze, et al.
Published: (2025)
by: Feng, Ze, et al.
Published: (2025)
Visual Question Answering Instruction: Unlocking Multimodal Large Language Model To Domain-Specific Visual Multitasks
by: Lee, Jusung, et al.
Published: (2024)
by: Lee, Jusung, et al.
Published: (2024)
Similar Items
-
Accelerating Pre-training of Multimodal LLMs via Chain-of-Sight
by: Huang, Ziyuan, et al.
Published: (2024) -
WSI-VQA: Interpreting Whole Slide Images by Generative Visual Question Answering
by: Chen, Pingyi, et al.
Published: (2024) -
PathVQ: Reforming Computational Pathology Foundation Model for Whole Slide Image Analysis via Vector Quantization
by: Li, Honglin, et al.
Published: (2025) -
Rethinking Transformer for Long Contextual Histopathology Whole Slide Image Analysis
by: Li, Honglin, et al.
Published: (2024) -
Attention-Challenging Multiple Instance Learning for Whole Slide Image Classification
by: Zhang, Yunlong, et al.
Published: (2023)