Learning How To Ask: Cycle-Consistency Refines Prompts in Multimodal Foundation Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Diesendruck, Maurice, Lin, Jianzhe, Imani, Shima, Mahalingam, Gayathri, Xu, Mingyang, Zhao, Jie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Diffusion-Augmented Coreset Expansion for Scalable Dataset Distillation
von: Abbasi, Ali, et al.
Veröffentlicht: (2024)
von: Abbasi, Ali, et al.
Veröffentlicht: (2024)
BatchPrompt: Accomplish more with less
von: Lin, Jianzhe, et al.
Veröffentlicht: (2023)
von: Lin, Jianzhe, et al.
Veröffentlicht: (2023)
Next-Token Prediction Task Assumes Optimal Data Ordering for LLM Training in Proof Generation
von: An, Chenyang, et al.
Veröffentlicht: (2024)
von: An, Chenyang, et al.
Veröffentlicht: (2024)
Multimodal Prompt Learning with Missing Modalities for Sentiment Analysis and Emotion Recognition
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
Dual-branch Prompting for Multimodal Machine Translation
von: Wang, Jie, et al.
Veröffentlicht: (2025)
von: Wang, Jie, et al.
Veröffentlicht: (2025)
How Easy is It to Fool Your Multimodal LLMs? An Empirical Analysis on Deceptive Prompts
von: Qian, Yusu, et al.
Veröffentlicht: (2024)
von: Qian, Yusu, et al.
Veröffentlicht: (2024)
Toward Robust Multimodal Learning using Multimodal Foundational Models
von: Zhao, Xianbing, et al.
Veröffentlicht: (2024)
von: Zhao, Xianbing, et al.
Veröffentlicht: (2024)
Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models
von: Zhou, Qiji, et al.
Veröffentlicht: (2024)
von: Zhou, Qiji, et al.
Veröffentlicht: (2024)
CycleChart: A Unified Consistency-Based Learning Framework for Bidirectional Chart Understanding and Generation
von: Deng, Dazhen, et al.
Veröffentlicht: (2025)
von: Deng, Dazhen, et al.
Veröffentlicht: (2025)
A Survey of Deep Learning for Geometry Problem Solving
von: Ma, Jianzhe, et al.
Veröffentlicht: (2025)
von: Ma, Jianzhe, et al.
Veröffentlicht: (2025)
Improving Text-to-Image Consistency via Automatic Prompt Optimization
von: Mañas, Oscar, et al.
Veröffentlicht: (2024)
von: Mañas, Oscar, et al.
Veröffentlicht: (2024)
RefereeBench: Are Video MLLMs Ready to be Multi-Sport Referees
von: Xu, Yichen, et al.
Veröffentlicht: (2026)
von: Xu, Yichen, et al.
Veröffentlicht: (2026)
Teaching Vision-Language Models to Ask: Resolving Ambiguity in Visual Questions
von: Jian, Pu, et al.
Veröffentlicht: (2025)
von: Jian, Pu, et al.
Veröffentlicht: (2025)
Toward Robust Incomplete Multimodal Sentiment Analysis via Hierarchical Representation Learning
von: Li, Mingcheng, et al.
Veröffentlicht: (2024)
von: Li, Mingcheng, et al.
Veröffentlicht: (2024)
MC-MKE: A Fine-Grained Multimodal Knowledge Editing Benchmark Emphasizing Modality Consistency
von: Zhang, Junzhe, et al.
Veröffentlicht: (2024)
von: Zhang, Junzhe, et al.
Veröffentlicht: (2024)
Promptception: How Sensitive Are Large Multimodal Models to Prompts?
von: Ismithdeen, Mohamed Insaf, et al.
Veröffentlicht: (2025)
von: Ismithdeen, Mohamed Insaf, et al.
Veröffentlicht: (2025)
Xuanwu: Evolving General Multimodal Models into an Industrial-Grade Foundation for Content Ecosystems
von: Zhang, Zhiqian, et al.
Veröffentlicht: (2026)
von: Zhang, Zhiqian, et al.
Veröffentlicht: (2026)
CPJ: Explainable Agricultural Pest Diagnosis via Caption-Prompt-Judge with LLM-Judged Refinement
von: Zhang, Wentao, et al.
Veröffentlicht: (2025)
von: Zhang, Wentao, et al.
Veröffentlicht: (2025)
TextTIGER: Text-based Intelligent Generation with Entity Prompt Refinement for Text-to-Image Generation
von: Ozaki, Shintaro, et al.
Veröffentlicht: (2025)
von: Ozaki, Shintaro, et al.
Veröffentlicht: (2025)
Ask Questions with Double Hints: Visual Question Generation with Answer-awareness and Region-reference
von: Shen, Kai, et al.
Veröffentlicht: (2024)
von: Shen, Kai, et al.
Veröffentlicht: (2024)
Visual Question Decomposition on Multimodal Large Language Models
von: Zhang, Haowei, et al.
Veröffentlicht: (2024)
von: Zhang, Haowei, et al.
Veröffentlicht: (2024)
Culture-Aware Humorous Captioning: Multimodal Humor Generation across Cultural Contexts
von: Xu, Run, et al.
Veröffentlicht: (2026)
von: Xu, Run, et al.
Veröffentlicht: (2026)
Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning
von: LASA Team, et al.
Veröffentlicht: (2025)
von: LASA Team, et al.
Veröffentlicht: (2025)
VP-MEL: Visual Prompts Guided Multimodal Entity Linking
von: Mi, Hongze, et al.
Veröffentlicht: (2024)
von: Mi, Hongze, et al.
Veröffentlicht: (2024)
Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking QA over Scientific Papers
von: Zhao, Yilun, et al.
Veröffentlicht: (2025)
von: Zhao, Yilun, et al.
Veröffentlicht: (2025)
Intern-S1: A Scientific Multimodal Foundation Model
von: Bai, Lei, et al.
Veröffentlicht: (2025)
von: Bai, Lei, et al.
Veröffentlicht: (2025)
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles
von: Slyman, Eric, et al.
Veröffentlicht: (2025)
von: Slyman, Eric, et al.
Veröffentlicht: (2025)
PromptMRG: Diagnosis-Driven Prompts for Medical Report Generation
von: Jin, Haibo, et al.
Veröffentlicht: (2023)
von: Jin, Haibo, et al.
Veröffentlicht: (2023)
Imp: Highly Capable Large Multimodal Models for Mobile Devices
von: Shao, Zhenwei, et al.
Veröffentlicht: (2024)
von: Shao, Zhenwei, et al.
Veröffentlicht: (2024)
AviationLMM: A Large Multimodal Foundation Model for Civil Aviation
von: Li, Wenbin, et al.
Veröffentlicht: (2026)
von: Li, Wenbin, et al.
Veröffentlicht: (2026)
Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages
von: Hu, Jinyi, et al.
Veröffentlicht: (2023)
von: Hu, Jinyi, et al.
Veröffentlicht: (2023)
Evolver: Chain-of-Evolution Prompting to Boost Large Multimodal Models for Hateful Meme Detection
von: Huang, Jinfa, et al.
Veröffentlicht: (2024)
von: Huang, Jinfa, et al.
Veröffentlicht: (2024)
Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale
von: Zou, Yicheng, et al.
Veröffentlicht: (2026)
von: Zou, Yicheng, et al.
Veröffentlicht: (2026)
TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
von: Shangguan, Ziyao, et al.
Veröffentlicht: (2024)
von: Shangguan, Ziyao, et al.
Veröffentlicht: (2024)
VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model
von: Song, Shezheng, et al.
Veröffentlicht: (2023)
von: Song, Shezheng, et al.
Veröffentlicht: (2023)
Rethinking Visual Prompting for Multimodal Large Language Models with External Knowledge
von: Lin, Yuanze, et al.
Veröffentlicht: (2024)
von: Lin, Yuanze, et al.
Veröffentlicht: (2024)
Historical Test-time Prompt Tuning for Vision Foundation Models
von: Zhang, Jingyi, et al.
Veröffentlicht: (2024)
von: Zhang, Jingyi, et al.
Veröffentlicht: (2024)
What You See is What You Ask: Evaluating Audio Descriptions
von: Kala, Divy, et al.
Veröffentlicht: (2025)
von: Kala, Divy, et al.
Veröffentlicht: (2025)
Contextual Image Attack: How Visual Context Exposes Multimodal Safety Vulnerabilities
von: Xiong, Yuan, et al.
Veröffentlicht: (2025)
von: Xiong, Yuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Diffusion-Augmented Coreset Expansion for Scalable Dataset Distillation
von: Abbasi, Ali, et al.
Veröffentlicht: (2024) -
BatchPrompt: Accomplish more with less
von: Lin, Jianzhe, et al.
Veröffentlicht: (2023) -
Next-Token Prediction Task Assumes Optimal Data Ordering for LLM Training in Proof Generation
von: An, Chenyang, et al.
Veröffentlicht: (2024) -
Multimodal Prompt Learning with Missing Modalities for Sentiment Analysis and Emotion Recognition
von: Guo, Zirun, et al.
Veröffentlicht: (2024) -
Dual-branch Prompting for Multimodal Machine Translation
von: Wang, Jie, et al.
Veröffentlicht: (2025)