Vision-Driven Prompt Optimization for Large Language Models in Multimodal Generative Tasks
Fuente:
arXiv
Salvato in:
| Autori principali: | Franklin, Leo, Boonmee, Apiradee, Wongsuwan, Kritsada |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Consultation on Industrial Machine Faults with Large language Models
di: Boonmee, Apiradee, et al.
Pubblicazione: (2024)
di: Boonmee, Apiradee, et al.
Pubblicazione: (2024)
Revolutionizing Radiology Workflow with Factual and Efficient CXR Report Generation
di: Sukjai, Pimchanok, et al.
Pubblicazione: (2025)
di: Sukjai, Pimchanok, et al.
Pubblicazione: (2025)
Context-Aware Multi-Turn Visual-Textual Reasoning in LVLMs via Dynamic Memory and Adaptive Visual Guidance
di: Shen, Weijie, et al.
Pubblicazione: (2025)
di: Shen, Weijie, et al.
Pubblicazione: (2025)
Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment
di: Yan, Ziang, et al.
Pubblicazione: (2024)
di: Yan, Ziang, et al.
Pubblicazione: (2024)
Self-Rewarding Large Vision-Language Models for Optimizing Prompts in Text-to-Image Generation
di: Yang, Hongji, et al.
Pubblicazione: (2025)
di: Yang, Hongji, et al.
Pubblicazione: (2025)
Griffon-G: Bridging Vision-Language and Vision-Centric Tasks via Large Multimodal Models
di: Zhan, Yufei, et al.
Pubblicazione: (2024)
di: Zhan, Yufei, et al.
Pubblicazione: (2024)
VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
di: Wu, Jiannan, et al.
Pubblicazione: (2024)
di: Wu, Jiannan, et al.
Pubblicazione: (2024)
RoboLLM: Robotic Vision Tasks Grounded on Multimodal Large Language Models
di: Long, Zijun, et al.
Pubblicazione: (2023)
di: Long, Zijun, et al.
Pubblicazione: (2023)
Multimodal Large Language Models for Medical Report Generation via Customized Prompt Tuning
di: Li, Chunlei, et al.
Pubblicazione: (2025)
di: Li, Chunlei, et al.
Pubblicazione: (2025)
Generalizing Vision-Language Models with Dedicated Prompt Guidance
di: Li, Xinyao, et al.
Pubblicazione: (2025)
di: Li, Xinyao, et al.
Pubblicazione: (2025)
Quantized Prompt for Efficient Generalization of Vision-Language Models
di: Hao, Tianxiang, et al.
Pubblicazione: (2024)
di: Hao, Tianxiang, et al.
Pubblicazione: (2024)
FedMVP: Federated Multimodal Visual Prompt Tuning for Vision-Language Models
di: Singha, Mainak, et al.
Pubblicazione: (2025)
di: Singha, Mainak, et al.
Pubblicazione: (2025)
Concept-Guided Prompt Learning for Generalization in Vision-Language Models
di: Zhang, Yi, et al.
Pubblicazione: (2024)
di: Zhang, Yi, et al.
Pubblicazione: (2024)
Multi-Turn Adaptive Prompting Attack on Large Vision-Language Models
di: Choi, In Chong, et al.
Pubblicazione: (2026)
di: Choi, In Chong, et al.
Pubblicazione: (2026)
Enhancing Multimodal Large Language Models with Multi-instance Visual Prompt Generator for Visual Representation Enrichment
di: Zhong, Wenliang, et al.
Pubblicazione: (2024)
di: Zhong, Wenliang, et al.
Pubblicazione: (2024)
TaskCLIP: Extend Large Vision-Language Model for Task Oriented Object Detection
di: Chen, Hanning, et al.
Pubblicazione: (2024)
di: Chen, Hanning, et al.
Pubblicazione: (2024)
Vision-Centric Activation and Coordination for Multimodal Large Language Models
di: Wang, Yunnan, et al.
Pubblicazione: (2025)
di: Wang, Yunnan, et al.
Pubblicazione: (2025)
Optimization of Prompt Learning via Multi-Knowledge Representation for Vision-Language Models
di: Zhang, Enming, et al.
Pubblicazione: (2024)
di: Zhang, Enming, et al.
Pubblicazione: (2024)
FMBench: Benchmarking Fairness in Multimodal Large Language Models on Medical Tasks
di: Wu, Peiran, et al.
Pubblicazione: (2024)
di: Wu, Peiran, et al.
Pubblicazione: (2024)
Interpretable Bilingual Multimodal Large Language Model for Diverse Biomedical Tasks
di: Wang, Lehan, et al.
Pubblicazione: (2024)
di: Wang, Lehan, et al.
Pubblicazione: (2024)
GeoRSMLLM: A Multimodal Large Language Model for Vision-Language Tasks in Geoscience and Remote Sensing
di: Zhang, Zilun, et al.
Pubblicazione: (2025)
di: Zhang, Zilun, et al.
Pubblicazione: (2025)
FashionLOGO: Prompting Multimodal Large Language Models for Fashion Logo Embeddings
di: Wang, Zhen, et al.
Pubblicazione: (2023)
di: Wang, Zhen, et al.
Pubblicazione: (2023)
Modeling Variants of Prompts for Vision-Language Models
di: Li, Ao, et al.
Pubblicazione: (2025)
di: Li, Ao, et al.
Pubblicazione: (2025)
Player-Centric Multimodal Prompt Generation for Large Language Model Based Identity-Aware Basketball Video Captioning
di: Xi, Zeyu, et al.
Pubblicazione: (2025)
di: Xi, Zeyu, et al.
Pubblicazione: (2025)
Effectively Enhancing Vision Language Large Models by Prompt Augmentation and Caption Utilization
di: Zhao, Minyi, et al.
Pubblicazione: (2024)
di: Zhao, Minyi, et al.
Pubblicazione: (2024)
Towards Language-Driven Video Inpainting via Multimodal Large Language Models
di: Wu, Jianzong, et al.
Pubblicazione: (2024)
di: Wu, Jianzong, et al.
Pubblicazione: (2024)
PromptKD: Unsupervised Prompt Distillation for Vision-Language Models
di: Li, Zheng, et al.
Pubblicazione: (2024)
di: Li, Zheng, et al.
Pubblicazione: (2024)
Advancing Multimodal In-Context Learning in Large Vision-Language Models with Task-aware Demonstrations
di: Li, Yanshu
Pubblicazione: (2025)
di: Li, Yanshu
Pubblicazione: (2025)
Prompting Large Vision-Language Models for Compositional Reasoning
di: Ossowski, Timothy, et al.
Pubblicazione: (2024)
di: Ossowski, Timothy, et al.
Pubblicazione: (2024)
Attention Prompting on Image for Large Vision-Language Models
di: Yu, Runpeng, et al.
Pubblicazione: (2024)
di: Yu, Runpeng, et al.
Pubblicazione: (2024)
GLOV: Guided Large Language Models as Implicit Optimizers for Vision Language Models
di: Mirza, M. Jehanzeb, et al.
Pubblicazione: (2024)
di: Mirza, M. Jehanzeb, et al.
Pubblicazione: (2024)
MAO: Efficient Model-Agnostic Optimization of Prompt Tuning for Vision-Language Models
di: Li, Haoyang, et al.
Pubblicazione: (2025)
di: Li, Haoyang, et al.
Pubblicazione: (2025)
LMOD: A Large Multimodal Ophthalmology Dataset and Benchmark for Large Vision-Language Models
di: Qin, Zhenyue, et al.
Pubblicazione: (2024)
di: Qin, Zhenyue, et al.
Pubblicazione: (2024)
InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models
di: Wang, Haomin, et al.
Pubblicazione: (2025)
di: Wang, Haomin, et al.
Pubblicazione: (2025)
VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language Models
di: Xu, Mingjie, et al.
Pubblicazione: (2025)
di: Xu, Mingjie, et al.
Pubblicazione: (2025)
ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language Models
di: Wu, Mingrui, et al.
Pubblicazione: (2024)
di: Wu, Mingrui, et al.
Pubblicazione: (2024)
Turbo: Informativity-Driven Acceleration Plug-In for Vision-Language Large Models
di: Ju, Chen, et al.
Pubblicazione: (2024)
di: Ju, Chen, et al.
Pubblicazione: (2024)
Vision-Language Models for Vision Tasks: A Survey
di: Zhang, Jingyi, et al.
Pubblicazione: (2023)
di: Zhang, Jingyi, et al.
Pubblicazione: (2023)
Active Prompt Learning in Vision Language Models
di: Bang, Jihwan, et al.
Pubblicazione: (2023)
di: Bang, Jihwan, et al.
Pubblicazione: (2023)
In the Era of Prompt Learning with Vision-Language Models
di: Jha, Ankit
Pubblicazione: (2024)
di: Jha, Ankit
Pubblicazione: (2024)
Documenti analoghi
-
Consultation on Industrial Machine Faults with Large language Models
di: Boonmee, Apiradee, et al.
Pubblicazione: (2024) -
Revolutionizing Radiology Workflow with Factual and Efficient CXR Report Generation
di: Sukjai, Pimchanok, et al.
Pubblicazione: (2025) -
Context-Aware Multi-Turn Visual-Textual Reasoning in LVLMs via Dynamic Memory and Adaptive Visual Guidance
di: Shen, Weijie, et al.
Pubblicazione: (2025) -
Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment
di: Yan, Ziang, et al.
Pubblicazione: (2024) -
Self-Rewarding Large Vision-Language Models for Optimizing Prompts in Text-to-Image Generation
di: Yang, Hongji, et al.
Pubblicazione: (2025)