DOS: Directional Object Separation in Text Embeddings for Multi-Object Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Byun, Dongnam, Park, Jungwon, Ko, Jungmin, Choi, Changin, Rhee, Wonjong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cross-Attention Head Position Patterns Can Align with Human Visual Concepts in Text-to-Image Generative Models
by: Park, Jungwon, et al.
Published: (2024)
by: Park, Jungwon, et al.
Published: (2024)
Selective Aggregation of Attention Maps Improves Diffusion-Based Visual Interpretation
by: Park, Jungwon, et al.
Published: (2026)
by: Park, Jungwon, et al.
Published: (2026)
Orthogonal Negative Guidance in Attention Feature Space for Text-to-Image Generation
by: Ko, Jungmin, et al.
Published: (2026)
by: Ko, Jungmin, et al.
Published: (2026)
Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering
by: Choi, Changin, et al.
Published: (2025)
by: Choi, Changin, et al.
Published: (2025)
Selectively Informative Description can Reduce Undesired Embedding Entanglements in Text-to-Image Personalization
by: Kim, Jimyeong, et al.
Published: (2024)
by: Kim, Jimyeong, et al.
Published: (2024)
An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM
by: Kim, Wonkyun, et al.
Published: (2024)
by: Kim, Wonkyun, et al.
Published: (2024)
ReFlex: Text-Guided Editing of Real Images in Rectified Flow via Mid-Step Feature Extraction and Attention Adaptation
by: Kim, Jimyeong, et al.
Published: (2025)
by: Kim, Jimyeong, et al.
Published: (2025)
Soft Head Selection for Injecting ICL-Derived Task Embeddings
by: Park, Jungwon, et al.
Published: (2025)
by: Park, Jungwon, et al.
Published: (2025)
Harmonizing Visual and Textual Embeddings for Zero-Shot Text-to-Image Customization
by: Song, Yeji, et al.
Published: (2024)
by: Song, Yeji, et al.
Published: (2024)
Task-Specific Preconditioner for Cross-Domain Few-Shot Learning
by: Kang, Suhyun, et al.
Published: (2024)
by: Kang, Suhyun, et al.
Published: (2024)
On-Off Pattern Encoding and Path-Count Encoding as Deep Neural Network Representations
by: Jung, Euna, et al.
Published: (2024)
by: Jung, Euna, et al.
Published: (2024)
Detecting Unknown Objects via Energy-based Separation for Open World Object Detection
by: Heo, Jun-Woo, et al.
Published: (2026)
by: Heo, Jun-Woo, et al.
Published: (2026)
InsideOut: Integrated RGB-Radiative Gaussian Splatting for Comprehensive 3D Object Representation
by: Lee, Jungmin, et al.
Published: (2025)
by: Lee, Jungmin, et al.
Published: (2025)
Compass Control: Multi Object Orientation Control for Text-to-Image Generation
by: Parihar, Rishubh, et al.
Published: (2025)
by: Parihar, Rishubh, et al.
Published: (2025)
Selectively Dilated Convolution for Accuracy-Preserving Sparse Pillar-based Embedded 3D Object Detection
by: Park, Seongmin, et al.
Published: (2024)
by: Park, Seongmin, et al.
Published: (2024)
When Confidence Misleads: Suffix Anchoring and Anchor-Proximity Confidence Modulation for Diffusion Language Models
by: Park, Jungwon, et al.
Published: (2026)
by: Park, Jungwon, et al.
Published: (2026)
Customizing Text-to-Image Diffusion with Object Viewpoint Control
by: Kumari, Nupur, et al.
Published: (2024)
by: Kumari, Nupur, et al.
Published: (2024)
Towards a Better Evaluation of Out-of-Domain Generalization
by: Hwang, Duhun, et al.
Published: (2024)
by: Hwang, Duhun, et al.
Published: (2024)
Object-Driven One-Shot Fine-tuning of Text-to-Image Diffusion with Prototypical Embedding
by: Lu, Jianxiang, et al.
Published: (2024)
by: Lu, Jianxiang, et al.
Published: (2024)
Improving Forward Compatibility in Class Incremental Learning by Increasing Representation Rank and Feature Richness
by: Kim, Jaeill, et al.
Published: (2024)
by: Kim, Jaeill, et al.
Published: (2024)
YOLO-Count: Differentiable Object Counting for Text-to-Image Generation
by: Zeng, Guanning, et al.
Published: (2025)
by: Zeng, Guanning, et al.
Published: (2025)
ComposeAnything: Composite Object Priors for Text-to-Image Generation
by: Khan, Zeeshan, et al.
Published: (2025)
by: Khan, Zeeshan, et al.
Published: (2025)
Object-Attribute Binding in Text-to-Image Generation: Evaluation and Control
by: Trusca, Maria Mihaela, et al.
Published: (2024)
by: Trusca, Maria Mihaela, et al.
Published: (2024)
From Detection to Association: Learning Discriminative Object Embeddings for Multi-Object Tracking
by: Shao, Yuqing, et al.
Published: (2025)
by: Shao, Yuqing, et al.
Published: (2025)
Under One Sun: Multi-Object Generative Perception of Materials and Illumination
by: Yoshii, Nobuo, et al.
Published: (2026)
by: Yoshii, Nobuo, et al.
Published: (2026)
Dual-Modality Anchor-Guided Filtering for Test-time Prompt Tuning
by: Choi, Jungwon, et al.
Published: (2026)
by: Choi, Jungwon, et al.
Published: (2026)
Captioning for Text-Video Retrieval via Dual-Group Direct Preference Optimization
by: Lee, Ji Soo, et al.
Published: (2025)
by: Lee, Ji Soo, et al.
Published: (2025)
Leveraging Image Augmentation for Object Manipulation: Towards Interpretable Controllability in Object-Centric Learning
by: Kim, Jinwoo, et al.
Published: (2023)
by: Kim, Jinwoo, et al.
Published: (2023)
OPEN: Object-wise Position Embedding for Multi-view 3D Object Detection
by: Hou, Jinghua, et al.
Published: (2024)
by: Hou, Jinghua, et al.
Published: (2024)
MOVi: Training-free Text-conditioned Multi-Object Video Generation
by: Rahman, Aimon, et al.
Published: (2025)
by: Rahman, Aimon, et al.
Published: (2025)
NCAP: Scene Text Image Super-Resolution with Non-CAtegorical Prior
by: Park, Dongwoo, et al.
Published: (2025)
by: Park, Dongwoo, et al.
Published: (2025)
OA-DET3D: Embedding Object Awareness as a General Plug-in for Multi-Camera 3D Object Detection
by: Chu, Xiaomeng, et al.
Published: (2023)
by: Chu, Xiaomeng, et al.
Published: (2023)
OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation
by: Oh, Yoonjin, et al.
Published: (2025)
by: Oh, Yoonjin, et al.
Published: (2025)
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking
by: Xuan, Shiyu, et al.
Published: (2025)
by: Xuan, Shiyu, et al.
Published: (2025)
Composing Object Relations and Attributes for Image-Text Matching
by: Pham, Khoi, et al.
Published: (2024)
by: Pham, Khoi, et al.
Published: (2024)
Evaluating Feature Attribution Methods for Electrocardiogram
by: Suh, Jangwon, et al.
Published: (2022)
by: Suh, Jangwon, et al.
Published: (2022)
M-PhyGs: Multi-Material Object Dynamics from Video
by: Wada, Norika, et al.
Published: (2025)
by: Wada, Norika, et al.
Published: (2025)
Is a Pure Transformer Effective for Separated and Online Multi-Object Tracking?
by: Liu, Chongwei, et al.
Published: (2024)
by: Liu, Chongwei, et al.
Published: (2024)
Multitwine: Multi-Object Compositing with Text and Layout Control
by: Tarrés, Gemma Canet, et al.
Published: (2025)
by: Tarrés, Gemma Canet, et al.
Published: (2025)
ST-LDM: A Universal Framework for Text-Grounded Object Generation in Real Images
by: Xue, Xiangtian, et al.
Published: (2024)
by: Xue, Xiangtian, et al.
Published: (2024)
Similar Items
-
Cross-Attention Head Position Patterns Can Align with Human Visual Concepts in Text-to-Image Generative Models
by: Park, Jungwon, et al.
Published: (2024) -
Selective Aggregation of Attention Maps Improves Diffusion-Based Visual Interpretation
by: Park, Jungwon, et al.
Published: (2026) -
Orthogonal Negative Guidance in Attention Feature Space for Text-to-Image Generation
by: Ko, Jungmin, et al.
Published: (2026) -
Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering
by: Choi, Changin, et al.
Published: (2025) -
Selectively Informative Description can Reduce Undesired Embedding Entanglements in Text-to-Image Personalization
by: Kim, Jimyeong, et al.
Published: (2024)