LLM as a Complementary Optimizer to Gradient Descent: A Case Study in Prompt Tuning
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Zixian, Liu, Ming, Ji, Zhilong, Bai, Jinfeng, Guo, Yiwen, Zuo, Wangmeng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MasterWeaver: Taming Editability and Face Identity for Personalized Text-to-Image Generation
by: Wei, Yuxiang, et al.
Published: (2024)
by: Wei, Yuxiang, et al.
Published: (2024)
MDIQA: Unified Image Quality Assessment for Multi-dimensional Evaluation and Restoration
by: Yao, Shunyu, et al.
Published: (2025)
by: Yao, Shunyu, et al.
Published: (2025)
VitaGlyph: Vitalizing Artistic Typography with Flexible Dual-branch Diffusion Models
by: Feng, Kailai, et al.
Published: (2024)
by: Feng, Kailai, et al.
Published: (2024)
Personalized Image Generation with Deep Generative Models: A Decade Survey
by: Wei, Yuxiang, et al.
Published: (2025)
by: Wei, Yuxiang, et al.
Published: (2025)
Spectral Prompt Tuning:Unveiling Unseen Classes for Zero-Shot Semantic Segmentation
by: Xu, Wenhao, et al.
Published: (2023)
by: Xu, Wenhao, et al.
Published: (2023)
FILP-3D: Enhancing 3D Few-shot Class-incremental Learning with Pre-trained Vision-Language Models
by: Xu, Wan, et al.
Published: (2023)
by: Xu, Wan, et al.
Published: (2023)
Dual Tuning for Reasoning Efficacy-Driven Data Curation in Multimodal LLM Training
by: Zheng, Ruobing, et al.
Published: (2026)
by: Zheng, Ruobing, et al.
Published: (2026)
TuneVLSeg: Prompt Tuning Benchmark for Vision-Language Segmentation Models
by: Adhikari, Rabin, et al.
Published: (2024)
by: Adhikari, Rabin, et al.
Published: (2024)
OPERA: An Agent for Image Restoration with End-to-End Joint Planning-Execution Optimization
by: Zhu, Feng, et al.
Published: (2026)
by: Zhu, Feng, et al.
Published: (2026)
Myriad: Large Multimodal Model by Applying Vision Experts for Industrial Anomaly Detection
by: Li, Yuanze, et al.
Published: (2023)
by: Li, Yuanze, et al.
Published: (2023)
Multimodal Prompt Learning with Missing Modalities for Sentiment Analysis and Emotion Recognition
by: Guo, Zirun, et al.
Published: (2024)
by: Guo, Zirun, et al.
Published: (2024)
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
by: Guo, Jarvis, et al.
Published: (2024)
by: Guo, Jarvis, et al.
Published: (2024)
HPNet: Dynamic Trajectory Forecasting with Historical Prediction Attention
by: Tang, Xiaolong, et al.
Published: (2024)
by: Tang, Xiaolong, et al.
Published: (2024)
Prompt-aligned Gradient for Prompt Tuning
by: Zhu, Beier, et al.
Published: (2022)
by: Zhu, Beier, et al.
Published: (2022)
The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generation
by: Gao, Bingjie, et al.
Published: (2025)
by: Gao, Bingjie, et al.
Published: (2025)
Patch-Prompt Aligned Bayesian Prompt Tuning for Vision-Language Models
by: Liu, Xinyang, et al.
Published: (2023)
by: Liu, Xinyang, et al.
Published: (2023)
Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM
by: Ji, Yatai, et al.
Published: (2024)
by: Ji, Yatai, et al.
Published: (2024)
Improving Transferability of Adversarial Examples via Bayesian Attacks
by: Li, Qizhang, et al.
Published: (2023)
by: Li, Qizhang, et al.
Published: (2023)
Optimizing Prompts for Text-to-Image Generation
by: Hao, Yaru, et al.
Published: (2022)
by: Hao, Yaru, et al.
Published: (2022)
Efficient Universal Goal Hijacking with Semantics-guided Prompt Organization
by: Huang, Yihao, et al.
Published: (2024)
by: Huang, Yihao, et al.
Published: (2024)
GLaVE-Cap: Global-Local Aligned Video Captioning with Vision Expert Integration
by: Xu, Wan, et al.
Published: (2025)
by: Xu, Wan, et al.
Published: (2025)
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
by: Du, Yifan, et al.
Published: (2023)
by: Du, Yifan, et al.
Published: (2023)
How to Utilize Complementary Vision-Text Information for 2D Structure Understanding
by: Dong, Jiancheng, et al.
Published: (2026)
by: Dong, Jiancheng, et al.
Published: (2026)
Prophet: Prompting Large Language Models with Complementary Answer Heuristics for Knowledge-based Visual Question Answering
by: Yu, Zhou, et al.
Published: (2023)
by: Yu, Zhou, et al.
Published: (2023)
Expert Pyramid Tuning: Efficient Parameter Fine-Tuning for Expertise-Driven Task Allocation
by: Zhang, Jia-Chen, et al.
Published: (2026)
by: Zhang, Jia-Chen, et al.
Published: (2026)
Ex-Omni: Enabling 3D Facial Animation Generation for Omni-modal Large Language Models
by: Zhang, Haoyu, et al.
Published: (2026)
by: Zhang, Haoyu, et al.
Published: (2026)
Universal Prompt Optimizer for Safe Text-to-Image Generation
by: Wu, Zongyu, et al.
Published: (2024)
by: Wu, Zongyu, et al.
Published: (2024)
Mitigating Object Hallucination via Robust Local Perception Search
by: Gao, Zixian, et al.
Published: (2025)
by: Gao, Zixian, et al.
Published: (2025)
m&m's: A Benchmark to Evaluate Tool-Use for multi-step multi-modal Tasks
by: Ma, Zixian, et al.
Published: (2024)
by: Ma, Zixian, et al.
Published: (2024)
Prompting4Debugging: Red-Teaming Text-to-Image Diffusion Models by Finding Problematic Prompts
by: Chin, Zhi-Yi, et al.
Published: (2023)
by: Chin, Zhi-Yi, et al.
Published: (2023)
Improving Text-to-Image Consistency via Automatic Prompt Optimization
by: Mañas, Oscar, et al.
Published: (2024)
by: Mañas, Oscar, et al.
Published: (2024)
Less is More: High-value Data Selection for Visual Instruction Tuning
by: Liu, Zikang, et al.
Published: (2024)
by: Liu, Zikang, et al.
Published: (2024)
Dual-branch Prompting for Multimodal Machine Translation
by: Wang, Jie, et al.
Published: (2025)
by: Wang, Jie, et al.
Published: (2025)
Explicit Relational Reasoning Network for Scene Text Detection
by: Su, Yuchen, et al.
Published: (2024)
by: Su, Yuchen, et al.
Published: (2024)
MuDPT: Multi-modal Deep-symphysis Prompt Tuning for Large Pre-trained Vision-Language Models
by: Miao, Yongzhu, et al.
Published: (2023)
by: Miao, Yongzhu, et al.
Published: (2023)
LayoutLLM: Layout Instruction Tuning with Large Language Models for Document Understanding
by: Luo, Chuwei, et al.
Published: (2024)
by: Luo, Chuwei, et al.
Published: (2024)
AnchorOPT: Towards Optimizing Dynamic Anchors for Adaptive Prompt Learning
by: Li, Zheng, et al.
Published: (2025)
by: Li, Zheng, et al.
Published: (2025)
CPJ: Explainable Agricultural Pest Diagnosis via Caption-Prompt-Judge with LLM-Judged Refinement
by: Zhang, Wentao, et al.
Published: (2025)
by: Zhang, Wentao, et al.
Published: (2025)
Classifier-guided Gradient Modulation for Enhanced Multimodal Learning
by: Guo, Zirun, et al.
Published: (2024)
by: Guo, Zirun, et al.
Published: (2024)
Historical Test-time Prompt Tuning for Vision Foundation Models
by: Zhang, Jingyi, et al.
Published: (2024)
by: Zhang, Jingyi, et al.
Published: (2024)
Similar Items
-
MasterWeaver: Taming Editability and Face Identity for Personalized Text-to-Image Generation
by: Wei, Yuxiang, et al.
Published: (2024) -
MDIQA: Unified Image Quality Assessment for Multi-dimensional Evaluation and Restoration
by: Yao, Shunyu, et al.
Published: (2025) -
VitaGlyph: Vitalizing Artistic Typography with Flexible Dual-branch Diffusion Models
by: Feng, Kailai, et al.
Published: (2024) -
Personalized Image Generation with Deep Generative Models: A Decade Survey
by: Wei, Yuxiang, et al.
Published: (2025) -
Spectral Prompt Tuning:Unveiling Unseen Classes for Zero-Shot Semantic Segmentation
by: Xu, Wenhao, et al.
Published: (2023)