Harmonizing Visual and Textual Embeddings for Zero-Shot Text-to-Image Customization
Fuente:
arXiv
Salvato in:
| Autori principali: | Song, Yeji, Kim, Jimyeong, Park, Wonhark, Shin, Wonsik, Rhee, Wonjong, Kwak, Nojun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ReFlex: Text-Guided Editing of Real Images in Rectified Flow via Mid-Step Feature Extraction and Attention Adaptation
di: Kim, Jimyeong, et al.
Pubblicazione: (2025)
di: Kim, Jimyeong, et al.
Pubblicazione: (2025)
Selectively Informative Description can Reduce Undesired Embedding Entanglements in Text-to-Image Personalization
di: Kim, Jimyeong, et al.
Pubblicazione: (2024)
di: Kim, Jimyeong, et al.
Pubblicazione: (2024)
Style Composition within Distinct LoRA modules for Traditional Art
di: Lee, Jaehyun, et al.
Pubblicazione: (2025)
di: Lee, Jaehyun, et al.
Pubblicazione: (2025)
Orthogonal Negative Guidance in Attention Feature Space for Text-to-Image Generation
di: Ko, Jungmin, et al.
Pubblicazione: (2026)
di: Ko, Jungmin, et al.
Pubblicazione: (2026)
Evaluating Feature Attribution Methods for Electrocardiogram
di: Suh, Jangwon, et al.
Pubblicazione: (2022)
di: Suh, Jangwon, et al.
Pubblicazione: (2022)
DOS: Directional Object Separation in Text Embeddings for Multi-Object Image Generation
di: Byun, Dongnam, et al.
Pubblicazione: (2025)
di: Byun, Dongnam, et al.
Pubblicazione: (2025)
Towards a Better Evaluation of Out-of-Domain Generalization
di: Hwang, Duhun, et al.
Pubblicazione: (2024)
di: Hwang, Duhun, et al.
Pubblicazione: (2024)
LoGoColor: Local-Global 3D Colorization for 360° Scenes
di: Chang, Yeonjin, et al.
Pubblicazione: (2025)
di: Chang, Yeonjin, et al.
Pubblicazione: (2025)
4DGS360: 360° Gaussian Reconstruction of Dynamic Objects from a Single Video
di: Jang, Jae Won, et al.
Pubblicazione: (2026)
di: Jang, Jae Won, et al.
Pubblicazione: (2026)
Enhancing Contrastive Learning with Efficient Combinatorial Positive Pairing
di: Kim, Jaeill, et al.
Pubblicazione: (2024)
di: Kim, Jaeill, et al.
Pubblicazione: (2024)
S3D: Sketch-Driven 3D Model Generation
di: Song, Hail, et al.
Pubblicazione: (2025)
di: Song, Hail, et al.
Pubblicazione: (2025)
Cross-Attention Head Position Patterns Can Align with Human Visual Concepts in Text-to-Image Generative Models
di: Park, Jungwon, et al.
Pubblicazione: (2024)
di: Park, Jungwon, et al.
Pubblicazione: (2024)
When Confidence Misleads: Suffix Anchoring and Anchor-Proximity Confidence Modulation for Diffusion Language Models
di: Park, Jungwon, et al.
Pubblicazione: (2026)
di: Park, Jungwon, et al.
Pubblicazione: (2026)
MSG Score: Automated Video Verification for Reliable Multi-Scene Generation
di: Yoon, Daewon, et al.
Pubblicazione: (2024)
di: Yoon, Daewon, et al.
Pubblicazione: (2024)
Point-to-Point: Sparse Motion Guidance for Controllable Video Editing
di: Song, Yeji, et al.
Pubblicazione: (2025)
di: Song, Yeji, et al.
Pubblicazione: (2025)
Conservative Generator, Progressive Discriminator: Coordination of Adversaries in Few-shot Incremental Image Synthesis
di: Kong, Chaerin, et al.
Pubblicazione: (2022)
di: Kong, Chaerin, et al.
Pubblicazione: (2022)
Selective Aggregation of Attention Maps Improves Diffusion-Based Visual Interpretation
di: Park, Jungwon, et al.
Pubblicazione: (2026)
di: Park, Jungwon, et al.
Pubblicazione: (2026)
DECOR:Decomposition and Projection of Text Embeddings for Text-to-Image Customization
di: Jang, Geonhui, et al.
Pubblicazione: (2024)
di: Jang, Geonhui, et al.
Pubblicazione: (2024)
Task-Specific Preconditioner for Cross-Domain Few-Shot Learning
di: Kang, Suhyun, et al.
Pubblicazione: (2024)
di: Kang, Suhyun, et al.
Pubblicazione: (2024)
An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM
di: Kim, Wonkyun, et al.
Pubblicazione: (2024)
di: Kim, Wonkyun, et al.
Pubblicazione: (2024)
Direct Consistency Optimization for Robust Customization of Text-to-Image Diffusion Models
di: Lee, Kyungmin, et al.
Pubblicazione: (2024)
di: Lee, Kyungmin, et al.
Pubblicazione: (2024)
MERLIN: Multimodal Embedding Refinement via LLM-based Iterative Navigation for Text-Video Retrieval-Rerank Pipeline
di: Han, Donghoon, et al.
Pubblicazione: (2024)
di: Han, Donghoon, et al.
Pubblicazione: (2024)
From Local to Global to Mechanistic: An iERF-Centered Unified Framework for Interpreting Vision Models
di: Kim, Yearim, et al.
Pubblicazione: (2026)
di: Kim, Yearim, et al.
Pubblicazione: (2026)
CSF: Black-box Fingerprinting via Compositional Semantics for Text-to-Image Models
di: Lee, Junhoo, et al.
Pubblicazione: (2026)
di: Lee, Junhoo, et al.
Pubblicazione: (2026)
Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering
di: Choi, Changin, et al.
Pubblicazione: (2025)
di: Choi, Changin, et al.
Pubblicazione: (2025)
Bi-ICE: An Inner Interpretable Framework for Image Classification via Bi-directional Interactions between Concept and Input Embeddings
di: Hong, Jinyung, et al.
Pubblicazione: (2024)
di: Hong, Jinyung, et al.
Pubblicazione: (2024)
DivCon-NeRF: Diverse and Consistent Ray Augmentation for Few-Shot NeRF
di: Lee, Ingyun, et al.
Pubblicazione: (2025)
di: Lee, Ingyun, et al.
Pubblicazione: (2025)
Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image Generator
di: Shin, Chaehun, et al.
Pubblicazione: (2024)
di: Shin, Chaehun, et al.
Pubblicazione: (2024)
Equilibrated Diffusion: Frequency-aware Textual Embedding for Equilibrated Image Customization
di: Ma, Liyuan, et al.
Pubblicazione: (2026)
di: Ma, Liyuan, et al.
Pubblicazione: (2026)
CustomText: Customized Textual Image Generation using Diffusion Models
di: Paliwal, Shubham, et al.
Pubblicazione: (2024)
di: Paliwal, Shubham, et al.
Pubblicazione: (2024)
Respect the model: Fine-grained and Robust Explanation with Sharing Ratio Decomposition
di: Han, Sangyu, et al.
Pubblicazione: (2024)
di: Han, Sangyu, et al.
Pubblicazione: (2024)
Causal Interpretation of Sparse Autoencoder Features in Vision
di: Han, Sangyu, et al.
Pubblicazione: (2025)
di: Han, Sangyu, et al.
Pubblicazione: (2025)
Improving Forward Compatibility in Class Incremental Learning by Increasing Representation Rank and Feature Richness
di: Kim, Jaeill, et al.
Pubblicazione: (2024)
di: Kim, Jaeill, et al.
Pubblicazione: (2024)
Zero-Shot Image Harmonization with Generative Model Prior
di: Chen, Jianqi, et al.
Pubblicazione: (2023)
di: Chen, Jianqi, et al.
Pubblicazione: (2023)
Multi-dimensional Preference Alignment by Conditioning Reward Itself
di: Jang, Jiho, et al.
Pubblicazione: (2025)
di: Jang, Jiho, et al.
Pubblicazione: (2025)
Fine-grained Textual Inversion Network for Zero-Shot Composed Image Retrieval
di: Lin, Haoqiang, et al.
Pubblicazione: (2025)
di: Lin, Haoqiang, et al.
Pubblicazione: (2025)
Extending CLIP's Image-Text Alignment to Referring Image Segmentation
di: Kim, Seoyeon, et al.
Pubblicazione: (2023)
di: Kim, Seoyeon, et al.
Pubblicazione: (2023)
The Role of Teacher Calibration in Knowledge Distillation
di: Kim, Suyoung, et al.
Pubblicazione: (2025)
di: Kim, Suyoung, et al.
Pubblicazione: (2025)
FontAdapter: Instant Font Adaptation in Visual Text Generation
di: Koo, Myungkyu, et al.
Pubblicazione: (2025)
di: Koo, Myungkyu, et al.
Pubblicazione: (2025)
Soft Head Selection for Injecting ICL-Derived Task Embeddings
di: Park, Jungwon, et al.
Pubblicazione: (2025)
di: Park, Jungwon, et al.
Pubblicazione: (2025)
Documenti analoghi
-
ReFlex: Text-Guided Editing of Real Images in Rectified Flow via Mid-Step Feature Extraction and Attention Adaptation
di: Kim, Jimyeong, et al.
Pubblicazione: (2025) -
Selectively Informative Description can Reduce Undesired Embedding Entanglements in Text-to-Image Personalization
di: Kim, Jimyeong, et al.
Pubblicazione: (2024) -
Style Composition within Distinct LoRA modules for Traditional Art
di: Lee, Jaehyun, et al.
Pubblicazione: (2025) -
Orthogonal Negative Guidance in Attention Feature Space for Text-to-Image Generation
di: Ko, Jungmin, et al.
Pubblicazione: (2026) -
Evaluating Feature Attribution Methods for Electrocardiogram
di: Suh, Jangwon, et al.
Pubblicazione: (2022)