InstanceDiffusion: Instance-level Control for Image Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Xudong, Darrell, Trevor, Rambhatla, Sai Saketh, Girdhar, Rohit, Misra, Ishan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Diffusion Autoencoders are Scalable Image Tokenizers
von: Chen, Yinbo, et al.
Veröffentlicht: (2025)
von: Chen, Yinbo, et al.
Veröffentlicht: (2025)
SelfEval: Leveraging the discriminative nature of generative models for evaluation
von: Rambhatla, Sai Saketh, et al.
Veröffentlicht: (2023)
von: Rambhatla, Sai Saketh, et al.
Veröffentlicht: (2023)
Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning
von: Girdhar, Rohit, et al.
Veröffentlicht: (2023)
von: Girdhar, Rohit, et al.
Veröffentlicht: (2023)
Generating Illustrated Instructions
von: Menon, Sachit, et al.
Veröffentlicht: (2023)
von: Menon, Sachit, et al.
Veröffentlicht: (2023)
MotiF: Making Text Count in Image Animation with Motion Focal Loss
von: Wang, Shijie, et al.
Veröffentlicht: (2024)
von: Wang, Shijie, et al.
Veröffentlicht: (2024)
Toward Diffusible High-Dimensional Latent Spaces: A Frequency Perspective
von: Lai, Bolin, et al.
Veröffentlicht: (2025)
von: Lai, Bolin, et al.
Veröffentlicht: (2025)
LLMs can see and hear without any training
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2025)
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2025)
The effectiveness of MAE pre-pretraining for billion-scale pretraining
von: Singh, Mannat, et al.
Veröffentlicht: (2023)
von: Singh, Mannat, et al.
Veröffentlicht: (2023)
MosaicFusion: Diffusion Models as Data Augmenters for Large Vocabulary Instance Segmentation
von: Xie, Jiahao, et al.
Veröffentlicht: (2023)
von: Xie, Jiahao, et al.
Veröffentlicht: (2023)
SC-MIL: Sparsely Coded Multiple Instance Learning for Whole Slide Image Classification
von: Qiu, Peijie, et al.
Veröffentlicht: (2023)
von: Qiu, Peijie, et al.
Veröffentlicht: (2023)
Occlusion-Ordered Semantic Instance Segmentation
von: Baselizadeh, Soroosh, et al.
Veröffentlicht: (2025)
von: Baselizadeh, Soroosh, et al.
Veröffentlicht: (2025)
Benchmarking the Robustness of Instance Segmentation Models
von: Dalva, Yusuf, et al.
Veröffentlicht: (2021)
von: Dalva, Yusuf, et al.
Veröffentlicht: (2021)
Continual Hyperbolic Learning of Instances and Classes
von: Ayoughi, Melika, et al.
Veröffentlicht: (2025)
von: Ayoughi, Melika, et al.
Veröffentlicht: (2025)
Lifting Embodied World Models for Planning and Control
von: Wang, Alex N., et al.
Veröffentlicht: (2026)
von: Wang, Alex N., et al.
Veröffentlicht: (2026)
PAIR-Diffusion: A Comprehensive Multimodal Object-Level Image Editor
von: Goel, Vidit, et al.
Veröffentlicht: (2023)
von: Goel, Vidit, et al.
Veröffentlicht: (2023)
Visual Lexicon: Rich Image Features in Language Space
von: Wang, XuDong, et al.
Veröffentlicht: (2024)
von: Wang, XuDong, et al.
Veröffentlicht: (2024)
SetFlow: Generating Structured Sets of Representations for Multiple Instance Learning
von: Jovišić, Nikola, et al.
Veröffentlicht: (2026)
von: Jovišić, Nikola, et al.
Veröffentlicht: (2026)
UnSAMv2: Self-Supervised Learning Enables Segment Anything at Any Granularity
von: Yu, Junwei, et al.
Veröffentlicht: (2025)
von: Yu, Junwei, et al.
Veröffentlicht: (2025)
UVIS: Unsupervised Video Instance Segmentation
von: Huang, Shuaiyi, et al.
Veröffentlicht: (2024)
von: Huang, Shuaiyi, et al.
Veröffentlicht: (2024)
Open-world Instance Segmentation: Top-down Learning with Bottom-up Supervision
von: Kalluri, Tarun, et al.
Veröffentlicht: (2023)
von: Kalluri, Tarun, et al.
Veröffentlicht: (2023)
Reconstruction Alignment Improves Unified Multimodal Models
von: Xie, Ji, et al.
Veröffentlicht: (2025)
von: Xie, Ji, et al.
Veröffentlicht: (2025)
3D-Aware Instance Segmentation and Tracking in Egocentric Videos
von: Bhalgat, Yash, et al.
Veröffentlicht: (2024)
von: Bhalgat, Yash, et al.
Veröffentlicht: (2024)
Hidden in plain sight: VLMs overlook their visual representations
von: Fu, Stephanie, et al.
Veröffentlicht: (2025)
von: Fu, Stephanie, et al.
Veröffentlicht: (2025)
Assessing SAM for Tree Crown Instance Segmentation from Drone Imagery
von: Teng, Mélisande, et al.
Veröffentlicht: (2025)
von: Teng, Mélisande, et al.
Veröffentlicht: (2025)
From Semantic To Instance: A Semi-Self-Supervised Learning Approach
von: Najafian, Keyhan, et al.
Veröffentlicht: (2025)
von: Najafian, Keyhan, et al.
Veröffentlicht: (2025)
Feature-Based Instance Neighbor Discovery: Advanced Stable Test-Time Adaptation in Dynamic World
von: Jiang, Qinting, et al.
Veröffentlicht: (2025)
von: Jiang, Qinting, et al.
Veröffentlicht: (2025)
Shape-Guided Diffusion with Inside-Outside Attention
von: Park, Dong Huk, et al.
Veröffentlicht: (2022)
von: Park, Dong Huk, et al.
Veröffentlicht: (2022)
IFAdapter: Instance Feature Control for Grounded Text-to-Image Generation
von: Wu, Yinwei, et al.
Veröffentlicht: (2024)
von: Wu, Yinwei, et al.
Veröffentlicht: (2024)
Face Density as a Proxy for Data Complexity: Quantifying the Hardness of Instance Count
von: Mohammadi-Seif, Abolfazl, et al.
Veröffentlicht: (2026)
von: Mohammadi-Seif, Abolfazl, et al.
Veröffentlicht: (2026)
Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction Detection
von: Seo, Soo Won, et al.
Veröffentlicht: (2026)
von: Seo, Soo Won, et al.
Veröffentlicht: (2026)
Prompting Foundation Models for Zero-Shot Ship Instance Segmentation in SAR Imagery
von: Mansour, Islam, et al.
Veröffentlicht: (2026)
von: Mansour, Islam, et al.
Veröffentlicht: (2026)
AROID: Improving Adversarial Robustness Through Online Instance-Wise Data Augmentation
von: Li, Lin, et al.
Veröffentlicht: (2023)
von: Li, Lin, et al.
Veröffentlicht: (2023)
SoftPQ: Robust Instance Segmentation Evaluation via Soft Matching and Tunable Thresholds
von: Karmakar, Ranit, et al.
Veröffentlicht: (2025)
von: Karmakar, Ranit, et al.
Veröffentlicht: (2025)
Low-Resolution Face Recognition via Adaptable Instance-Relation Distillation
von: Shi, Ruixin, et al.
Veröffentlicht: (2024)
von: Shi, Ruixin, et al.
Veröffentlicht: (2024)
REOrdering Patches Improves Vision Models
von: Kutscher, Declan, et al.
Veröffentlicht: (2025)
von: Kutscher, Declan, et al.
Veröffentlicht: (2025)
Foveated Instance Segmentation
von: Zeng, Hongyi, et al.
Veröffentlicht: (2025)
von: Zeng, Hongyi, et al.
Veröffentlicht: (2025)
How Effective Can Dropout Be in Multiple Instance Learning ?
von: Zhu, Wenhui, et al.
Veröffentlicht: (2025)
von: Zhu, Wenhui, et al.
Veröffentlicht: (2025)
Diffuse and Disperse: Image Generation with Representation Regularization
von: Wang, Runqian, et al.
Veröffentlicht: (2025)
von: Wang, Runqian, et al.
Veröffentlicht: (2025)
Compositional Chain-of-Thought Prompting for Large Multimodal Models
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
von: Qin, Yiming, et al.
Veröffentlicht: (2025)
von: Qin, Yiming, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Diffusion Autoencoders are Scalable Image Tokenizers
von: Chen, Yinbo, et al.
Veröffentlicht: (2025) -
SelfEval: Leveraging the discriminative nature of generative models for evaluation
von: Rambhatla, Sai Saketh, et al.
Veröffentlicht: (2023) -
Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning
von: Girdhar, Rohit, et al.
Veröffentlicht: (2023) -
Generating Illustrated Instructions
von: Menon, Sachit, et al.
Veröffentlicht: (2023) -
MotiF: Making Text Count in Image Animation with Motion Focal Loss
von: Wang, Shijie, et al.
Veröffentlicht: (2024)