Textual and Visual Prompt Fusion for Image Editing via Step-Wise Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Feng, Zhanbo, Ling, Zenan, Lu, Xinyu, Gong, Ci, Zhou, Feng, Bao, Wugedele, Li, Jie, Yang, Fan, Qiu, Robert C. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Series-to-Series Diffusion Bridge Model
von: Yang, Hao, et al.
Veröffentlicht: (2024)
von: Yang, Hao, et al.
Veröffentlicht: (2024)
Textualize Visual Prompt for Image Editing via Diffusion Bridge
von: Xu, Pengcheng, et al.
Veröffentlicht: (2025)
von: Xu, Pengcheng, et al.
Veröffentlicht: (2025)
Deep Equilibrium Models are Almost Equivalent to Not-so-deep Explicit Models for High-dimensional Gaussian Mixtures
von: Ling, Zenan, et al.
Veröffentlicht: (2024)
von: Ling, Zenan, et al.
Veröffentlicht: (2024)
Adaptive Discretization for Consistency Models
von: Bai, Jiayu, et al.
Veröffentlicht: (2025)
von: Bai, Jiayu, et al.
Veröffentlicht: (2025)
IGNN-Solver: A Graph Neural Solver for Implicit Graph Neural Networks
von: Lin, Junchao, et al.
Veröffentlicht: (2024)
von: Lin, Junchao, et al.
Veröffentlicht: (2024)
Visual Textualization for Image Prompted Object Detection
von: Wu, Yongjian, et al.
Veröffentlicht: (2025)
von: Wu, Yongjian, et al.
Veröffentlicht: (2025)
Dreamer: Dual-RIS-aided Imager in Complementary Modes
von: Wang, Fuhai, et al.
Veröffentlicht: (2024)
von: Wang, Fuhai, et al.
Veröffentlicht: (2024)
Robust and Communication-Efficient Federated Domain Adaptation via Random Features
von: Feng, Zhanbo, et al.
Veröffentlicht: (2023)
von: Feng, Zhanbo, et al.
Veröffentlicht: (2023)
Continual Learning on CLIP via Incremental Prompt Tuning with Intrinsic Textual Anchors
von: Lu, Haodong, et al.
Veröffentlicht: (2025)
von: Lu, Haodong, et al.
Veröffentlicht: (2025)
BeDAViN Sound Events Dataset for Audio-Visual Embodied Navigation
von: Shi, Zhanbo
Veröffentlicht: (2025)
von: Shi, Zhanbo
Veröffentlicht: (2025)
An Item is Worth a Prompt: Versatile Image Editing with Disentangled Control
von: Feng, Aosong, et al.
Veröffentlicht: (2024)
von: Feng, Aosong, et al.
Veröffentlicht: (2024)
DeepDefense: Layer-Wise Gradient-Feature Alignment for Building Robust Neural Networks
von: Lin, Ci, et al.
Veröffentlicht: (2025)
von: Lin, Ci, et al.
Veröffentlicht: (2025)
PromptHub: Enhancing Multi-Prompt Visual In-Context Learning with Locality-Aware Fusion, Concentration and Alignment
von: Luo, Tianci, et al.
Veröffentlicht: (2026)
von: Luo, Tianci, et al.
Veröffentlicht: (2026)
The Fusion of Infrared and Visible Images via Feature Extraction and Subwindow Variance Filtering
von: Xin Feng, et al.
Veröffentlicht: (2024)
von: Xin Feng, et al.
Veröffentlicht: (2024)
Deep Learning‐Based Building Change Detection in Off‐Nadir Images via a Pixel‐Wise and Patch‐Wise Fusion Strategy
von: Jianfeng Huang, et al.
Veröffentlicht: (2025)
von: Jianfeng Huang, et al.
Veröffentlicht: (2025)
Beyond Textual CoT: Interleaved Text-Image Chains with Deep Confidence Reasoning for Image Editing
von: Zou, Zhentao, et al.
Veröffentlicht: (2025)
von: Zou, Zhentao, et al.
Veröffentlicht: (2025)
Nonstationary Sparse Spectral Permanental Process
von: Sun, Zicheng, et al.
Veröffentlicht: (2024)
von: Sun, Zicheng, et al.
Veröffentlicht: (2024)
WiseEdit: Benchmarking Cognition- and Creativity-Informed Image Editing
von: Pan, Kaihang, et al.
Veröffentlicht: (2025)
von: Pan, Kaihang, et al.
Veröffentlicht: (2025)
Revisiting Logistic-softmax Likelihood in Bayesian Meta-Learning for Few-Shot Classification
von: Ke, Tianjun, et al.
Veröffentlicht: (2023)
von: Ke, Tianjun, et al.
Veröffentlicht: (2023)
Distilling Textual Priors from LLM to Efficient Image Fusion
von: Zhang, Ran, et al.
Veröffentlicht: (2025)
von: Zhang, Ran, et al.
Veröffentlicht: (2025)
Step-Wise Formal Verification for LLM-Based Mathematical Problem Solving
von: Zhou, Kuo, et al.
Veröffentlicht: (2025)
von: Zhou, Kuo, et al.
Veröffentlicht: (2025)
Beyond Images: Adaptive Fusion of Visual and Textual Data for Food Classification
von: Mittal, Prateek, et al.
Veröffentlicht: (2023)
von: Mittal, Prateek, et al.
Veröffentlicht: (2023)
COCA: Classifier-Oriented Calibration via Textual Prototype for Source-Free Universal Domain Adaptation
von: Liu, Xinghong, et al.
Veröffentlicht: (2023)
von: Liu, Xinghong, et al.
Veröffentlicht: (2023)
ViTA-PAR: Visual and Textual Attribute Alignment with Attribute Prompting for Pedestrian Attribute Recognition
von: Park, Minjeong, et al.
Veröffentlicht: (2025)
von: Park, Minjeong, et al.
Veröffentlicht: (2025)
SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs
von: Yin, Yuanyang, et al.
Veröffentlicht: (2024)
von: Yin, Yuanyang, et al.
Veröffentlicht: (2024)
Dragging with Geometry: From Pixels to Geometry-Guided Image Editing
von: Pu, Xinyu, et al.
Veröffentlicht: (2025)
von: Pu, Xinyu, et al.
Veröffentlicht: (2025)
Towards Visuospatial Cognition via Hierarchical Fusion of Visual Experts
von: Feng, Qi
Veröffentlicht: (2025)
von: Feng, Qi
Veröffentlicht: (2025)
Diving into Kronecker Adapters: Component Design Matters
von: Bai, Jiayu, et al.
Veröffentlicht: (2026)
von: Bai, Jiayu, et al.
Veröffentlicht: (2026)
FreeInpaint: Tuning-free Prompt Alignment and Visual Rationality Enhancement in Image Inpainting
von: Gong, Chao, et al.
Veröffentlicht: (2025)
von: Gong, Chao, et al.
Veröffentlicht: (2025)
Letter: Amino Acid Imbalance Is an Independent Factor for Mortality in Patients With Liver Cirrhosis
von: Zhanbo Qu
Veröffentlicht: (2026)
von: Zhanbo Qu
Veröffentlicht: (2026)
Codebook Configuration for RIS-aided Systems via Implicit Neural Representations
von: Yang, Huiying, et al.
Veröffentlicht: (2023)
von: Yang, Huiying, et al.
Veröffentlicht: (2023)
Consistency Deep Equilibrium Models
von: Lin, Junchao, et al.
Veröffentlicht: (2026)
von: Lin, Junchao, et al.
Veröffentlicht: (2026)
TIE: Revolutionizing Text-based Image Editing for Complex-Prompt Following and High-Fidelity Editing
von: Zhang, Xinyu, et al.
Veröffentlicht: (2024)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2024)
LocalAlign: Enabling Generalizable Prompt Injection Defense via Generation of Near-Target Adversarial Examples for Alignment Training
von: Gong, Yuyang, et al.
Veröffentlicht: (2026)
von: Gong, Yuyang, et al.
Veröffentlicht: (2026)
Cognitive-Inspired Hierarchical Attention Fusion With Visual and Textual for Cross-Domain Sequential Recommendation
von: Wu, Wangyu, et al.
Veröffentlicht: (2025)
von: Wu, Wangyu, et al.
Veröffentlicht: (2025)
Chain-of-Jailbreak Attack for Image Generation Models via Editing Step by Step
von: Wang, Wenxuan, et al.
Veröffentlicht: (2024)
von: Wang, Wenxuan, et al.
Veröffentlicht: (2024)
Empowering Visual Creativity: A Vision-Language Assistant to Image Editing Recommendations
von: Shen, Tiancheng, et al.
Veröffentlicht: (2024)
von: Shen, Tiancheng, et al.
Veröffentlicht: (2024)
Hyperbolic Cycle Alignment for Infrared-Visible Image Fusion
von: Li, Timing, et al.
Veröffentlicht: (2025)
von: Li, Timing, et al.
Veröffentlicht: (2025)
FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts
von: Gong, Yichen, et al.
Veröffentlicht: (2023)
von: Gong, Yichen, et al.
Veröffentlicht: (2023)
FG-CLIP: Fine-Grained Visual and Textual Alignment
von: Xie, Chunyu, et al.
Veröffentlicht: (2025)
von: Xie, Chunyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Series-to-Series Diffusion Bridge Model
von: Yang, Hao, et al.
Veröffentlicht: (2024) -
Textualize Visual Prompt for Image Editing via Diffusion Bridge
von: Xu, Pengcheng, et al.
Veröffentlicht: (2025) -
Deep Equilibrium Models are Almost Equivalent to Not-so-deep Explicit Models for High-dimensional Gaussian Mixtures
von: Ling, Zenan, et al.
Veröffentlicht: (2024) -
Adaptive Discretization for Consistency Models
von: Bai, Jiayu, et al.
Veröffentlicht: (2025) -
IGNN-Solver: A Graph Neural Solver for Implicit Graph Neural Networks
von: Lin, Junchao, et al.
Veröffentlicht: (2024)