CUPID: Contextual Understanding of Prompt-conditioned Image Distributions
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Yayan, Li, Mingwei, Berger, Matthew |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Graphical Perception of Saliency-based Model Explanations
by: Zhao, Yayan, et al.
Published: (2024)
by: Zhao, Yayan, et al.
Published: (2024)
Fluid Grey 2: How Well Does Generative Adversarial Network Learn Deeper Topology Structure in Architecture That Matches Images?
by: Qiu, Yayan, et al.
Published: (2025)
by: Qiu, Yayan, et al.
Published: (2025)
Image Generation from Contextually-Contradictory Prompts
by: Huberman, Saar, et al.
Published: (2025)
by: Huberman, Saar, et al.
Published: (2025)
ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding
by: Lee, Hosu, et al.
Published: (2025)
by: Lee, Hosu, et al.
Published: (2025)
Aquila-plus: Prompt-Driven Visual-Language Models for Pixel-Level Remote Sensing Image Understanding
by: Lu, Kaixuan
Published: (2024)
by: Lu, Kaixuan
Published: (2024)
Dual Prompting Image Restoration with Diffusion Transformers
by: Kong, Dehong, et al.
Published: (2025)
by: Kong, Dehong, et al.
Published: (2025)
ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models
by: Yi, Jingwei, et al.
Published: (2025)
by: Yi, Jingwei, et al.
Published: (2025)
Prompt Optimizer of Text-to-Image Diffusion Models for Abstract Concept Understanding
by: Fan, Zezhong, et al.
Published: (2024)
by: Fan, Zezhong, et al.
Published: (2024)
Spatial Understanding from Videos: Structured Prompts Meet Simulation Data
by: Zhang, Haoyu, et al.
Published: (2025)
by: Zhang, Haoyu, et al.
Published: (2025)
SOWing Information: Cultivating Contextual Coherence with MLLMs in Image Generation
by: Pei, Yuhan, et al.
Published: (2024)
by: Pei, Yuhan, et al.
Published: (2024)
Token-Efficient Multimodal Reasoning via Image Prompt Packaging
by: Choi, Joong Ho, et al.
Published: (2026)
by: Choi, Joong Ho, et al.
Published: (2026)
Prompt Decoupling for Text-to-Image Person Re-identification
by: Li, Weihao, et al.
Published: (2024)
by: Li, Weihao, et al.
Published: (2024)
Prompt-SID: Learning Structural Representation Prompt via Latent Diffusion for Single-Image Denoising
by: Li, Huaqiu, et al.
Published: (2025)
by: Li, Huaqiu, et al.
Published: (2025)
RieMind: Geometry-Grounded Spatial Agent for Scene Understanding
by: Ropero, Fernando, et al.
Published: (2026)
by: Ropero, Fernando, et al.
Published: (2026)
Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs
by: Wang, Haochen, et al.
Published: (2025)
by: Wang, Haochen, et al.
Published: (2025)
Hierarchical Prompt Learning for Image- and Text-Based Person Re-Identification
by: Zhou, Linhan, et al.
Published: (2025)
by: Zhou, Linhan, et al.
Published: (2025)
PTCMIL: Multiple Instance Learning via Prompt Token Clustering for Whole Slide Image Analysis
by: Zhao, Beidi, et al.
Published: (2025)
by: Zhao, Beidi, et al.
Published: (2025)
ArtAug: Enhancing Text-to-Image Generation through Synthesis-Understanding Interaction
by: Duan, Zhongjie, et al.
Published: (2024)
by: Duan, Zhongjie, et al.
Published: (2024)
Enhancing Multimodal Understanding with CLIP-Based Image-to-Text Transformation
by: Che, Chang, et al.
Published: (2024)
by: Che, Chang, et al.
Published: (2024)
Decoding the Pulse of Reasoning VLMs in Multi-Image Understanding Tasks
by: Li, Chenjun
Published: (2026)
by: Li, Chenjun
Published: (2026)
Embedded Visual Prompt Tuning
by: Zu, Wenqiang, et al.
Published: (2024)
by: Zu, Wenqiang, et al.
Published: (2024)
Face-MakeUp: Multimodal Facial Prompts for Text-to-Image Generation
by: Dai, Dawei, et al.
Published: (2025)
by: Dai, Dawei, et al.
Published: (2025)
Dynamic Prompt Optimizing for Text-to-Image Generation
by: Mo, Wenyi, et al.
Published: (2024)
by: Mo, Wenyi, et al.
Published: (2024)
TC-LLaVA: Rethinking the Transfer from Image to Video Understanding with Temporal Considerations
by: Gao, Mingze, et al.
Published: (2024)
by: Gao, Mingze, et al.
Published: (2024)
Understanding Pure Textual Reasoning for Blind Image Quality Assessment
by: Li, Yuan, et al.
Published: (2026)
by: Li, Yuan, et al.
Published: (2026)
Multimodal Prompt Decoupling Attack on the Safety Filters in Text-to-Image Models
by: Peng, Xingkai, et al.
Published: (2025)
by: Peng, Xingkai, et al.
Published: (2025)
Brain Tumor Segmentation in MRI Images with 3D U-Net and Contextual Transformer
by: Nguyen, Thien-Qua T., et al.
Published: (2024)
by: Nguyen, Thien-Qua T., et al.
Published: (2024)
Controllable Contextualized Image Captioning: Directing the Visual Narrative through User-Defined Highlights
by: Mao, Shunqi, et al.
Published: (2024)
by: Mao, Shunqi, et al.
Published: (2024)
PrismVAU: Prompt-Refined Inference System for Multimodal Video Anomaly Understanding
by: Erregue, Iñaki, et al.
Published: (2026)
by: Erregue, Iñaki, et al.
Published: (2026)
Evaluating Contextual Intelligence in Recyclability: A Comprehensive Study of Image-Based Reasoning Systems
by: Park, Eliot, et al.
Published: (2025)
by: Park, Eliot, et al.
Published: (2025)
Batch-Instructed Gradient for Prompt Evolution:Systematic Prompt Optimization for Enhanced Text-to-Image Synthesis
by: Yang, Xinrui, et al.
Published: (2024)
by: Yang, Xinrui, et al.
Published: (2024)
RS-GPT4V: A Unified Multimodal Instruction-Following Dataset for Remote Sensing Image Understanding
by: Xu, Linrui, et al.
Published: (2024)
by: Xu, Linrui, et al.
Published: (2024)
Towards Accurate UAV Image Perception: Guiding Vision-Language Models with Stronger Task Prompts
by: Guo, Mingning, et al.
Published: (2025)
by: Guo, Mingning, et al.
Published: (2025)
Any4D: Open-Prompt 4D Generation from Natural Language and Images
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
DetailMaster: Can Your Text-to-Image Model Handle Long Prompts?
by: Jiao, Qirui, et al.
Published: (2025)
by: Jiao, Qirui, et al.
Published: (2025)
Curriculum Prompting Foundation Models for Medical Image Segmentation
by: Zheng, Xiuqi, et al.
Published: (2024)
by: Zheng, Xiuqi, et al.
Published: (2024)
Patch-enhanced Mask Encoder Prompt Image Generation
by: Xu, Shusong, et al.
Published: (2024)
by: Xu, Shusong, et al.
Published: (2024)
Attention Prompting on Image for Large Vision-Language Models
by: Yu, Runpeng, et al.
Published: (2024)
by: Yu, Runpeng, et al.
Published: (2024)
CHEM: Estimating and Understanding Hallucinations in Deep Learning for Image Processing
by: Li, Jianfei, et al.
Published: (2025)
by: Li, Jianfei, et al.
Published: (2025)
Image Captions are Natural Prompts for Text-to-Image Models
by: Lei, Shiye, et al.
Published: (2023)
by: Lei, Shiye, et al.
Published: (2023)
Similar Items
-
Graphical Perception of Saliency-based Model Explanations
by: Zhao, Yayan, et al.
Published: (2024) -
Fluid Grey 2: How Well Does Generative Adversarial Network Learn Deeper Topology Structure in Architecture That Matches Images?
by: Qiu, Yayan, et al.
Published: (2025) -
Image Generation from Contextually-Contradictory Prompts
by: Huberman, Saar, et al.
Published: (2025) -
ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding
by: Lee, Hosu, et al.
Published: (2025) -
Aquila-plus: Prompt-Driven Visual-Language Models for Pixel-Level Remote Sensing Image Understanding
by: Lu, Kaixuan
Published: (2024)