CUPID: Contextual Understanding of Prompt-conditioned Image Distributions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Yayan, Li, Mingwei, Berger, Matthew |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Graphical Perception of Saliency-based Model Explanations
von: Zhao, Yayan, et al.
Veröffentlicht: (2024)
von: Zhao, Yayan, et al.
Veröffentlicht: (2024)
Fluid Grey 2: How Well Does Generative Adversarial Network Learn Deeper Topology Structure in Architecture That Matches Images?
von: Qiu, Yayan, et al.
Veröffentlicht: (2025)
von: Qiu, Yayan, et al.
Veröffentlicht: (2025)
Image Generation from Contextually-Contradictory Prompts
von: Huberman, Saar, et al.
Veröffentlicht: (2025)
von: Huberman, Saar, et al.
Veröffentlicht: (2025)
ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding
von: Lee, Hosu, et al.
Veröffentlicht: (2025)
von: Lee, Hosu, et al.
Veröffentlicht: (2025)
Aquila-plus: Prompt-Driven Visual-Language Models for Pixel-Level Remote Sensing Image Understanding
von: Lu, Kaixuan
Veröffentlicht: (2024)
von: Lu, Kaixuan
Veröffentlicht: (2024)
Dual Prompting Image Restoration with Diffusion Transformers
von: Kong, Dehong, et al.
Veröffentlicht: (2025)
von: Kong, Dehong, et al.
Veröffentlicht: (2025)
ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models
von: Yi, Jingwei, et al.
Veröffentlicht: (2025)
von: Yi, Jingwei, et al.
Veröffentlicht: (2025)
Prompt Optimizer of Text-to-Image Diffusion Models for Abstract Concept Understanding
von: Fan, Zezhong, et al.
Veröffentlicht: (2024)
von: Fan, Zezhong, et al.
Veröffentlicht: (2024)
Spatial Understanding from Videos: Structured Prompts Meet Simulation Data
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
SOWing Information: Cultivating Contextual Coherence with MLLMs in Image Generation
von: Pei, Yuhan, et al.
Veröffentlicht: (2024)
von: Pei, Yuhan, et al.
Veröffentlicht: (2024)
Token-Efficient Multimodal Reasoning via Image Prompt Packaging
von: Choi, Joong Ho, et al.
Veröffentlicht: (2026)
von: Choi, Joong Ho, et al.
Veröffentlicht: (2026)
Prompt Decoupling for Text-to-Image Person Re-identification
von: Li, Weihao, et al.
Veröffentlicht: (2024)
von: Li, Weihao, et al.
Veröffentlicht: (2024)
Prompt-SID: Learning Structural Representation Prompt via Latent Diffusion for Single-Image Denoising
von: Li, Huaqiu, et al.
Veröffentlicht: (2025)
von: Li, Huaqiu, et al.
Veröffentlicht: (2025)
RieMind: Geometry-Grounded Spatial Agent for Scene Understanding
von: Ropero, Fernando, et al.
Veröffentlicht: (2026)
von: Ropero, Fernando, et al.
Veröffentlicht: (2026)
Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
Hierarchical Prompt Learning for Image- and Text-Based Person Re-Identification
von: Zhou, Linhan, et al.
Veröffentlicht: (2025)
von: Zhou, Linhan, et al.
Veröffentlicht: (2025)
PTCMIL: Multiple Instance Learning via Prompt Token Clustering for Whole Slide Image Analysis
von: Zhao, Beidi, et al.
Veröffentlicht: (2025)
von: Zhao, Beidi, et al.
Veröffentlicht: (2025)
ArtAug: Enhancing Text-to-Image Generation through Synthesis-Understanding Interaction
von: Duan, Zhongjie, et al.
Veröffentlicht: (2024)
von: Duan, Zhongjie, et al.
Veröffentlicht: (2024)
Enhancing Multimodal Understanding with CLIP-Based Image-to-Text Transformation
von: Che, Chang, et al.
Veröffentlicht: (2024)
von: Che, Chang, et al.
Veröffentlicht: (2024)
Decoding the Pulse of Reasoning VLMs in Multi-Image Understanding Tasks
von: Li, Chenjun
Veröffentlicht: (2026)
von: Li, Chenjun
Veröffentlicht: (2026)
Embedded Visual Prompt Tuning
von: Zu, Wenqiang, et al.
Veröffentlicht: (2024)
von: Zu, Wenqiang, et al.
Veröffentlicht: (2024)
Face-MakeUp: Multimodal Facial Prompts for Text-to-Image Generation
von: Dai, Dawei, et al.
Veröffentlicht: (2025)
von: Dai, Dawei, et al.
Veröffentlicht: (2025)
Dynamic Prompt Optimizing for Text-to-Image Generation
von: Mo, Wenyi, et al.
Veröffentlicht: (2024)
von: Mo, Wenyi, et al.
Veröffentlicht: (2024)
TC-LLaVA: Rethinking the Transfer from Image to Video Understanding with Temporal Considerations
von: Gao, Mingze, et al.
Veröffentlicht: (2024)
von: Gao, Mingze, et al.
Veröffentlicht: (2024)
Understanding Pure Textual Reasoning for Blind Image Quality Assessment
von: Li, Yuan, et al.
Veröffentlicht: (2026)
von: Li, Yuan, et al.
Veröffentlicht: (2026)
Multimodal Prompt Decoupling Attack on the Safety Filters in Text-to-Image Models
von: Peng, Xingkai, et al.
Veröffentlicht: (2025)
von: Peng, Xingkai, et al.
Veröffentlicht: (2025)
Brain Tumor Segmentation in MRI Images with 3D U-Net and Contextual Transformer
von: Nguyen, Thien-Qua T., et al.
Veröffentlicht: (2024)
von: Nguyen, Thien-Qua T., et al.
Veröffentlicht: (2024)
Controllable Contextualized Image Captioning: Directing the Visual Narrative through User-Defined Highlights
von: Mao, Shunqi, et al.
Veröffentlicht: (2024)
von: Mao, Shunqi, et al.
Veröffentlicht: (2024)
PrismVAU: Prompt-Refined Inference System for Multimodal Video Anomaly Understanding
von: Erregue, Iñaki, et al.
Veröffentlicht: (2026)
von: Erregue, Iñaki, et al.
Veröffentlicht: (2026)
Evaluating Contextual Intelligence in Recyclability: A Comprehensive Study of Image-Based Reasoning Systems
von: Park, Eliot, et al.
Veröffentlicht: (2025)
von: Park, Eliot, et al.
Veröffentlicht: (2025)
Batch-Instructed Gradient for Prompt Evolution:Systematic Prompt Optimization for Enhanced Text-to-Image Synthesis
von: Yang, Xinrui, et al.
Veröffentlicht: (2024)
von: Yang, Xinrui, et al.
Veröffentlicht: (2024)
RS-GPT4V: A Unified Multimodal Instruction-Following Dataset for Remote Sensing Image Understanding
von: Xu, Linrui, et al.
Veröffentlicht: (2024)
von: Xu, Linrui, et al.
Veröffentlicht: (2024)
Towards Accurate UAV Image Perception: Guiding Vision-Language Models with Stronger Task Prompts
von: Guo, Mingning, et al.
Veröffentlicht: (2025)
von: Guo, Mingning, et al.
Veröffentlicht: (2025)
Any4D: Open-Prompt 4D Generation from Natural Language and Images
von: Li, Hao, et al.
Veröffentlicht: (2025)
von: Li, Hao, et al.
Veröffentlicht: (2025)
DetailMaster: Can Your Text-to-Image Model Handle Long Prompts?
von: Jiao, Qirui, et al.
Veröffentlicht: (2025)
von: Jiao, Qirui, et al.
Veröffentlicht: (2025)
Curriculum Prompting Foundation Models for Medical Image Segmentation
von: Zheng, Xiuqi, et al.
Veröffentlicht: (2024)
von: Zheng, Xiuqi, et al.
Veröffentlicht: (2024)
Patch-enhanced Mask Encoder Prompt Image Generation
von: Xu, Shusong, et al.
Veröffentlicht: (2024)
von: Xu, Shusong, et al.
Veröffentlicht: (2024)
Attention Prompting on Image for Large Vision-Language Models
von: Yu, Runpeng, et al.
Veröffentlicht: (2024)
von: Yu, Runpeng, et al.
Veröffentlicht: (2024)
CHEM: Estimating and Understanding Hallucinations in Deep Learning for Image Processing
von: Li, Jianfei, et al.
Veröffentlicht: (2025)
von: Li, Jianfei, et al.
Veröffentlicht: (2025)
Image Captions are Natural Prompts for Text-to-Image Models
von: Lei, Shiye, et al.
Veröffentlicht: (2023)
von: Lei, Shiye, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Graphical Perception of Saliency-based Model Explanations
von: Zhao, Yayan, et al.
Veröffentlicht: (2024) -
Fluid Grey 2: How Well Does Generative Adversarial Network Learn Deeper Topology Structure in Architecture That Matches Images?
von: Qiu, Yayan, et al.
Veröffentlicht: (2025) -
Image Generation from Contextually-Contradictory Prompts
von: Huberman, Saar, et al.
Veröffentlicht: (2025) -
ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding
von: Lee, Hosu, et al.
Veröffentlicht: (2025) -
Aquila-plus: Prompt-Driven Visual-Language Models for Pixel-Level Remote Sensing Image Understanding
von: Lu, Kaixuan
Veröffentlicht: (2024)