VisualPrompter: Semantic-Aware Prompt Optimization with Visual Feedback for Text-to-Image Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Shiyu, Sun, Mingzhen, Wang, Weining, Wang, Yequan, Liu, Jing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DiffPrompter: Differentiable Implicit Visual Prompts for Semantic-Segmentation in Adverse Conditions
by: Kalwar, Sanket, et al.
Published: (2023)
by: Kalwar, Sanket, et al.
Published: (2023)
OmniDFA: A Unified Framework for Open Set Synthesis Image Detection and Few-Shot Attribution
by: Wu, Shiyu, et al.
Published: (2025)
by: Wu, Shiyu, et al.
Published: (2025)
Few-Shot Learner Generalizes Across AI-Generated Image Detection
by: Wu, Shiyu, et al.
Published: (2025)
by: Wu, Shiyu, et al.
Published: (2025)
COMUNI: Decomposing Common and Unique Video Signals for Diffusion-based Video Generation
by: Sun, Mingzhen, et al.
Published: (2024)
by: Sun, Mingzhen, et al.
Published: (2024)
IP-Prompter: Training-Free Theme-Specific Image Generation via Dynamic Visual Prompting
by: Zhang, Yuxin, et al.
Published: (2025)
by: Zhang, Yuxin, et al.
Published: (2025)
AttriPrompter: Auto-Prompting with Attribute Semantics for Zero-shot Nuclei Detection via Visual-Language Pre-trained Models
by: Wu, Yongjian, et al.
Published: (2024)
by: Wu, Yongjian, et al.
Published: (2024)
LocateEdit-Bench: A Benchmark for Instruction-Based Editing Localization
by: Wu, Shiyu, et al.
Published: (2026)
by: Wu, Shiyu, et al.
Published: (2026)
MM-LDM: Multi-Modal Latent Diffusion Model for Sounding Video Generation
by: Sun, Mingzhen, et al.
Published: (2024)
by: Sun, Mingzhen, et al.
Published: (2024)
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
by: Wu, Yin, et al.
Published: (2025)
by: Wu, Yin, et al.
Published: (2025)
Semantic-Enriched Latent Visual Reasoning
by: Xu, Tianrun, et al.
Published: (2026)
by: Xu, Tianrun, et al.
Published: (2026)
Unbiased Object Detection Beyond Frequency with Visually Prompted Image Synthesis
by: Cai, Xinhao, et al.
Published: (2025)
by: Cai, Xinhao, et al.
Published: (2025)
Visual Prompt-Agnostic Evolution
by: Wang, Junze, et al.
Published: (2026)
by: Wang, Junze, et al.
Published: (2026)
Universal Prompt Optimizer for Safe Text-to-Image Generation
by: Wu, Zongyu, et al.
Published: (2024)
by: Wu, Zongyu, et al.
Published: (2024)
Noise Diffusion for Enhancing Semantic Faithfulness in Text-to-Image Synthesis
by: Miao, Boming, et al.
Published: (2024)
by: Miao, Boming, et al.
Published: (2024)
Learning Unknown Spoof Prompts for Generalized Face Anti-Spoofing Using Only Real Face Images
by: Jiang, Fangling, et al.
Published: (2025)
by: Jiang, Fangling, et al.
Published: (2025)
Mismatch Quest: Visual and Textual Feedback for Image-Text Misalignment
by: Gordon, Brian, et al.
Published: (2023)
by: Gordon, Brian, et al.
Published: (2023)
Panoptic Captioning: An Equivalence Bridge for Image and Text
by: Lin, Kun-Yu, et al.
Published: (2025)
by: Lin, Kun-Yu, et al.
Published: (2025)
Text-guided Visual Prompt DINO for Generic Segmentation
by: Guan, Yuchen, et al.
Published: (2025)
by: Guan, Yuchen, et al.
Published: (2025)
ProAPO: Progressively Automatic Prompt Optimization for Visual Classification
by: Qu, Xiangyan, et al.
Published: (2025)
by: Qu, Xiangyan, et al.
Published: (2025)
DTLLM-VLT: Diverse Text Generation for Visual Language Tracking Based on LLM
by: Li, Xuchen, et al.
Published: (2024)
by: Li, Xuchen, et al.
Published: (2024)
Verify Claimed Text-to-Image Models via Boundary-Aware Prompt Optimization
by: Zhao, Zidong, et al.
Published: (2026)
by: Zhao, Zidong, et al.
Published: (2026)
Emergent Visual-Semantic Hierarchies in Image-Text Representations
by: Alper, Morris, et al.
Published: (2024)
by: Alper, Morris, et al.
Published: (2024)
Visual Prompt Discovery via Semantic Exploration
by: Kim, Jaechang, et al.
Published: (2026)
by: Kim, Jaechang, et al.
Published: (2026)
PromptHub: Enhancing Multi-Prompt Visual In-Context Learning with Locality-Aware Fusion, Concentration and Alignment
by: Luo, Tianci, et al.
Published: (2026)
by: Luo, Tianci, et al.
Published: (2026)
Variation-Aware Semantic Image Synthesis
by: Xu, Mingle, et al.
Published: (2023)
by: Xu, Mingle, et al.
Published: (2023)
TIPO: Text to Image with Text Presampling for Prompt Optimization
by: Yeh, Shih-Ying, et al.
Published: (2024)
by: Yeh, Shih-Ying, et al.
Published: (2024)
UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and Perception
by: Song, Xinyang, et al.
Published: (2025)
by: Song, Xinyang, et al.
Published: (2025)
DA-VPT: Semantic-Guided Visual Prompt Tuning for Vision Transformers
by: Ren, Li, et al.
Published: (2025)
by: Ren, Li, et al.
Published: (2025)
Visual Textualization for Image Prompted Object Detection
by: Wu, Yongjian, et al.
Published: (2025)
by: Wu, Yongjian, et al.
Published: (2025)
Learning Visual Proxy for Compositional Zero-Shot Learning
by: Zhang, Shiyu, et al.
Published: (2025)
by: Zhang, Shiyu, et al.
Published: (2025)
Prompt-Softbox-Prompt: A Free-Text Embedding Control for Image Editing
by: Yang, Yitong, et al.
Published: (2024)
by: Yang, Yitong, et al.
Published: (2024)
HIPTrack: Visual Tracking with Historical Prompts
by: Cai, Wenrui, et al.
Published: (2023)
by: Cai, Wenrui, et al.
Published: (2023)
Batch-Instructed Gradient for Prompt Evolution:Systematic Prompt Optimization for Enhanced Text-to-Image Synthesis
by: Yang, Xinrui, et al.
Published: (2024)
by: Yang, Xinrui, et al.
Published: (2024)
Text-to-CAD Generation Through Infusing Visual Feedback in Large Language Models
by: Wang, Ruiyu, et al.
Published: (2025)
by: Wang, Ruiyu, et al.
Published: (2025)
TV-3DG: Mastering Text-to-3D Customized Generation with Visual Prompt
by: Yang, Jiahui, et al.
Published: (2024)
by: Yang, Jiahui, et al.
Published: (2024)
Towards Reliable and Holistic Visual In-Context Learning Prompt Selection
by: Wu, Wenxiao, et al.
Published: (2025)
by: Wu, Wenxiao, et al.
Published: (2025)
AR-Diffusion: Asynchronous Video Generation with Auto-Regressive Diffusion
by: Sun, Mingzhen, et al.
Published: (2025)
by: Sun, Mingzhen, et al.
Published: (2025)
RepText: Rendering Visual Text via Replicating
by: Wang, Haofan, et al.
Published: (2025)
by: Wang, Haofan, et al.
Published: (2025)
Visual Semantic Description Generation with MLLMs for Image-Text Matching
by: Chen, Junyu, et al.
Published: (2025)
by: Chen, Junyu, et al.
Published: (2025)
PhysCorr: Dual-Reward DPO for Physics-Constrained Text-to-Video Generation with Automated Preference Selection
by: Wang, Peiyao, et al.
Published: (2025)
by: Wang, Peiyao, et al.
Published: (2025)
Similar Items
-
DiffPrompter: Differentiable Implicit Visual Prompts for Semantic-Segmentation in Adverse Conditions
by: Kalwar, Sanket, et al.
Published: (2023) -
OmniDFA: A Unified Framework for Open Set Synthesis Image Detection and Few-Shot Attribution
by: Wu, Shiyu, et al.
Published: (2025) -
Few-Shot Learner Generalizes Across AI-Generated Image Detection
by: Wu, Shiyu, et al.
Published: (2025) -
COMUNI: Decomposing Common and Unique Video Signals for Diffusion-based Video Generation
by: Sun, Mingzhen, et al.
Published: (2024) -
IP-Prompter: Training-Free Theme-Specific Image Generation via Dynamic Visual Prompting
by: Zhang, Yuxin, et al.
Published: (2025)