Dynamic Prompting of Frozen Text-to-Image Diffusion Models for Panoptic Narrative Grounding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Hongyu, Hui, Tianrui, Ding, Zihan, Zhang, Jing, Ma, Bin, Wei, Xiaoming, Han, Jizhong, Liu, Si |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Enriching Phrases with Coupled Pixel and Object Contexts for Panoptic Narrative Grounding
von: Hui, Tianrui, et al.
Veröffentlicht: (2023)
von: Hui, Tianrui, et al.
Veröffentlicht: (2023)
Unleashing the Temporal-Spatial Reasoning Capacity of GPT for Training-Free Audio and Language Referenced Video Object Segmentation
von: Huang, Shaofei, et al.
Veröffentlicht: (2024)
von: Huang, Shaofei, et al.
Veröffentlicht: (2024)
Panoptic Segmentation of Mammograms with Text-To-Image Diffusion Model
von: Zhao, Kun, et al.
Veröffentlicht: (2024)
von: Zhao, Kun, et al.
Veröffentlicht: (2024)
LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding
von: Li, Hongyu, et al.
Veröffentlicht: (2025)
von: Li, Hongyu, et al.
Veröffentlicht: (2025)
Panoptic Captioning: An Equivalence Bridge for Image and Text
von: Lin, Kun-Yu, et al.
Veröffentlicht: (2025)
von: Lin, Kun-Yu, et al.
Veröffentlicht: (2025)
Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation
von: Zhang, Wenchao, et al.
Veröffentlicht: (2025)
von: Zhang, Wenchao, et al.
Veröffentlicht: (2025)
Unveiling Structural Memorization: Structural Membership Inference Attack for Text-to-Image Diffusion Models
von: Li, Qiao, et al.
Veröffentlicht: (2024)
von: Li, Qiao, et al.
Veröffentlicht: (2024)
EMMA: Your Text-to-Image Diffusion Model Can Secretly Accept Multi-Modal Prompts
von: Han, Yucheng, et al.
Veröffentlicht: (2024)
von: Han, Yucheng, et al.
Veröffentlicht: (2024)
FreeEdit: Mask-free Reference-based Image Editing with Multi-modal Instruction
von: He, Runze, et al.
Veröffentlicht: (2024)
von: He, Runze, et al.
Veröffentlicht: (2024)
GroundingBooth: Grounding Text-to-Image Customization
von: Xiong, Zhexiao, et al.
Veröffentlicht: (2024)
von: Xiong, Zhexiao, et al.
Veröffentlicht: (2024)
Exploring Phrase-Level Grounding with Text-to-Image Diffusion Model
von: Yang, Danni, et al.
Veröffentlicht: (2024)
von: Yang, Danni, et al.
Veröffentlicht: (2024)
EDITOR: Effective and Interpretable Prompt Inversion for Text-to-Image Diffusion Models
von: Li, Mingzhe, et al.
Veröffentlicht: (2025)
von: Li, Mingzhe, et al.
Veröffentlicht: (2025)
F-LMM: Grounding Frozen Large Multimodal Models
von: Wu, Size, et al.
Veröffentlicht: (2024)
von: Wu, Size, et al.
Veröffentlicht: (2024)
TextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models
von: Xiao, Yao, et al.
Veröffentlicht: (2025)
von: Xiao, Yao, et al.
Veröffentlicht: (2025)
Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation Models
von: Fu, Shenghao, et al.
Veröffentlicht: (2024)
von: Fu, Shenghao, et al.
Veröffentlicht: (2024)
PromptLA: Towards Integrity Verification of Black-box Text-to-Image Diffusion Models
von: Zhang, Zhuomeng, et al.
Veröffentlicht: (2024)
von: Zhang, Zhuomeng, et al.
Veröffentlicht: (2024)
TextPSG: Panoptic Scene Graph Generation from Textual Descriptions
von: Zhao, Chengyang, et al.
Veröffentlicht: (2023)
von: Zhao, Chengyang, et al.
Veröffentlicht: (2023)
Contrastive Prompts Improve Disentanglement in Text-to-Image Diffusion Models
von: Wu, Chen, et al.
Veröffentlicht: (2024)
von: Wu, Chen, et al.
Veröffentlicht: (2024)
Instant Preference Alignment for Text-to-Image Diffusion Models
von: Li, Yang, et al.
Veröffentlicht: (2025)
von: Li, Yang, et al.
Veröffentlicht: (2025)
GVDIFF: Grounded Text-to-Video Generation with Diffusion Models
von: Dou, Huanzhang, et al.
Veröffentlicht: (2024)
von: Dou, Huanzhang, et al.
Veröffentlicht: (2024)
Image Super-Resolution with Text Prompt Diffusion
von: Chen, Zheng, et al.
Veröffentlicht: (2023)
von: Chen, Zheng, et al.
Veröffentlicht: (2023)
Beyond Text: Frozen Large Language Models in Visual Signal Comprehension
von: Zhu, Lei, et al.
Veröffentlicht: (2024)
von: Zhu, Lei, et al.
Veröffentlicht: (2024)
PoGDiff: Product-of-Gaussians Diffusion Models for Imbalanced Text-to-Image Generation
von: Wang, Ziyan, et al.
Veröffentlicht: (2025)
von: Wang, Ziyan, et al.
Veröffentlicht: (2025)
Prompting4Debugging: Red-Teaming Text-to-Image Diffusion Models by Finding Problematic Prompts
von: Chin, Zhi-Yi, et al.
Veröffentlicht: (2023)
von: Chin, Zhi-Yi, et al.
Veröffentlicht: (2023)
Prompt-Free Conditional Diffusion for Multi-object Image Augmentation
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
Model Will Tell: Training Membership Inference for Diffusion Models
von: Fu, Xiaomeng, et al.
Veröffentlicht: (2024)
von: Fu, Xiaomeng, et al.
Veröffentlicht: (2024)
FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training
von: Cao, Anjia, et al.
Veröffentlicht: (2024)
von: Cao, Anjia, et al.
Veröffentlicht: (2024)
A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models
von: Shuai, Xincheng, et al.
Veröffentlicht: (2024)
von: Shuai, Xincheng, et al.
Veröffentlicht: (2024)
Supercharged One-step Text-to-Image Diffusion Models with Negative Prompts
von: Nguyen, Viet, et al.
Veröffentlicht: (2024)
von: Nguyen, Viet, et al.
Veröffentlicht: (2024)
Toward Early Quality Assessment of Text-to-Image Diffusion Models
von: Guo, Huanlei, et al.
Veröffentlicht: (2026)
von: Guo, Huanlei, et al.
Veröffentlicht: (2026)
Panoptic-FlashOcc: An Efficient Baseline to Marry Semantic Occupancy with Panoptic via Instance Center
von: Yu, Zichen, et al.
Veröffentlicht: (2024)
von: Yu, Zichen, et al.
Veröffentlicht: (2024)
LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models
von: Lian, Long, et al.
Veröffentlicht: (2023)
von: Lian, Long, et al.
Veröffentlicht: (2023)
VMix: Improving Text-to-Image Diffusion Model with Cross-Attention Mixing Control
von: Wu, Shaojin, et al.
Veröffentlicht: (2024)
von: Wu, Shaojin, et al.
Veröffentlicht: (2024)
VLPrompt: Vision-Language Prompting for Panoptic Scene Graph Generation
von: Zhou, Zijian, et al.
Veröffentlicht: (2023)
von: Zhou, Zijian, et al.
Veröffentlicht: (2023)
Grounding Text-to-Image Diffusion Models for Controlled High-Quality Image Generation
von: Süleyman, Ahmad, et al.
Veröffentlicht: (2025)
von: Süleyman, Ahmad, et al.
Veröffentlicht: (2025)
InitNO: Boosting Text-to-Image Diffusion Models via Initial Noise Optimization
von: Guo, Xiefan, et al.
Veröffentlicht: (2024)
von: Guo, Xiefan, et al.
Veröffentlicht: (2024)
Prompt Generation Networks for Input-Space Adaptation of Frozen Vision Transformers
von: Loedeman, Jochem, et al.
Veröffentlicht: (2022)
von: Loedeman, Jochem, et al.
Veröffentlicht: (2022)
Not All Prompts Are Made Equal: Prompt-based Pruning of Text-to-Image Diffusion Models
von: Ganjdanesh, Alireza, et al.
Veröffentlicht: (2024)
von: Ganjdanesh, Alireza, et al.
Veröffentlicht: (2024)
EmoGen: Emotional Image Content Generation with Text-to-Image Diffusion Models
von: Yang, Jingyuan, et al.
Veröffentlicht: (2024)
von: Yang, Jingyuan, et al.
Veröffentlicht: (2024)
Talk is Not Always Cheap: Promoting Wireless Sensing Models with Text Prompts
von: Yang, Zhenkui, et al.
Veröffentlicht: (2025)
von: Yang, Zhenkui, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Enriching Phrases with Coupled Pixel and Object Contexts for Panoptic Narrative Grounding
von: Hui, Tianrui, et al.
Veröffentlicht: (2023) -
Unleashing the Temporal-Spatial Reasoning Capacity of GPT for Training-Free Audio and Language Referenced Video Object Segmentation
von: Huang, Shaofei, et al.
Veröffentlicht: (2024) -
Panoptic Segmentation of Mammograms with Text-To-Image Diffusion Model
von: Zhao, Kun, et al.
Veröffentlicht: (2024) -
LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding
von: Li, Hongyu, et al.
Veröffentlicht: (2025) -
Panoptic Captioning: An Equivalence Bridge for Image and Text
von: Lin, Kun-Yu, et al.
Veröffentlicht: (2025)