ST-LDM: A Universal Framework for Text-Grounded Object Generation in Real Images
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xue, Xiangtian, Wu, Jiasong, Kong, Youyong, Senhadji, Lotfi, Shu, Huazhong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Rethinking Referring Object Removal
von: Xue, Xiangtian, et al.
Veröffentlicht: (2024)
von: Xue, Xiangtian, et al.
Veröffentlicht: (2024)
Multiscale Low-Frequency Memory Network for Improved Feature Extraction in Convolutional Neural Networks
von: Wu, Fuzhi, et al.
Veröffentlicht: (2024)
von: Wu, Fuzhi, et al.
Veröffentlicht: (2024)
GGT-100K: Generative Ground Truth for Generalizable Real-World Image Restoration
von: Kong, Xiangtao, et al.
Veröffentlicht: (2026)
von: Kong, Xiangtao, et al.
Veröffentlicht: (2026)
TiFRe: Text-guided Video Frame Reduction for Efficient Video Multi-modal Large Language Models
von: Zheng, Xiangtian, et al.
Veröffentlicht: (2026)
von: Zheng, Xiangtian, et al.
Veröffentlicht: (2026)
Intraoperative 2D/3D Registration via Spherical Similarity Learning and Differentiable Levenberg-Marquardt Optimization
von: Chen, Minheng, et al.
Veröffentlicht: (2025)
von: Chen, Minheng, et al.
Veröffentlicht: (2025)
ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning
von: Guo, Yuxiang, et al.
Veröffentlicht: (2025)
von: Guo, Yuxiang, et al.
Veröffentlicht: (2025)
LDM-Morph: Latent diffusion model guided deformable image registration
von: Wu, Jiong, et al.
Veröffentlicht: (2024)
von: Wu, Jiong, et al.
Veröffentlicht: (2024)
StructLDM: Structured Latent Diffusion for 3D Human Generation
von: Hu, Tao, et al.
Veröffentlicht: (2024)
von: Hu, Tao, et al.
Veröffentlicht: (2024)
Detecting AutoEncoder is Enough to Catch LDM Generated Images
von: Vesnin, Dmitry, et al.
Veröffentlicht: (2024)
von: Vesnin, Dmitry, et al.
Veröffentlicht: (2024)
MM-LDM: Multi-Modal Latent Diffusion Model for Sounding Video Generation
von: Sun, Mingzhen, et al.
Veröffentlicht: (2024)
von: Sun, Mingzhen, et al.
Veröffentlicht: (2024)
RT-OVAD: Real-Time Open-Vocabulary Aerial Object Detection via Image-Text Collaboration
von: Wei, Guoting, et al.
Veröffentlicht: (2024)
von: Wei, Guoting, et al.
Veröffentlicht: (2024)
RepLDM: Reprogramming Pretrained Latent Diffusion Models for High-Quality, High-Efficiency, High-Resolution Image Generation
von: Cao, Boyuan, et al.
Veröffentlicht: (2024)
von: Cao, Boyuan, et al.
Veröffentlicht: (2024)
GroundingBooth: Grounding Text-to-Image Customization
von: Xiong, Zhexiao, et al.
Veröffentlicht: (2024)
von: Xiong, Zhexiao, et al.
Veröffentlicht: (2024)
Real-time Video Target Tracking Algorithm Utilizing Convolutional Neural Networks (CNN)
von: Tan, Chaoyi, et al.
Veröffentlicht: (2024)
von: Tan, Chaoyi, et al.
Veröffentlicht: (2024)
SG-LDM: Semantic-Guided LiDAR Generation via Latent-Aligned Diffusion
von: Xiang, Zhengkang, et al.
Veröffentlicht: (2025)
von: Xiang, Zhengkang, et al.
Veröffentlicht: (2025)
DOS: Directional Object Separation in Text Embeddings for Multi-Object Image Generation
von: Byun, Dongnam, et al.
Veröffentlicht: (2025)
von: Byun, Dongnam, et al.
Veröffentlicht: (2025)
Universal Prompt Optimizer for Safe Text-to-Image Generation
von: Wu, Zongyu, et al.
Veröffentlicht: (2024)
von: Wu, Zongyu, et al.
Veröffentlicht: (2024)
FlashEval: Towards Fast and Accurate Evaluation of Text-to-image Diffusion Generative Models
von: Zhao, Lin, et al.
Veröffentlicht: (2024)
von: Zhao, Lin, et al.
Veröffentlicht: (2024)
PhyEdit: Towards Real-World Object Manipulation via Physically-Grounded Image Editing
von: Xu, Ruihang, et al.
Veröffentlicht: (2026)
von: Xu, Ruihang, et al.
Veröffentlicht: (2026)
DyST-XL: Dynamic Layout Planning and Content Control for Compositional Text-to-Video Generation
von: He, Weijie, et al.
Veröffentlicht: (2025)
von: He, Weijie, et al.
Veröffentlicht: (2025)
Target-aware Bidirectional Fusion Transformer for Aerial Object Tracking
von: Sun, Xinglong, et al.
Veröffentlicht: (2025)
von: Sun, Xinglong, et al.
Veröffentlicht: (2025)
R2LDM: An Efficient 4D Radar Super-Resolution Framework Leveraging Diffusion Model
von: Zheng, Boyuan, et al.
Veröffentlicht: (2025)
von: Zheng, Boyuan, et al.
Veröffentlicht: (2025)
HCMA: Hierarchical Cross-model Alignment for Grounded Text-to-Image Generation
von: Wang, Hang, et al.
Veröffentlicht: (2025)
von: Wang, Hang, et al.
Veröffentlicht: (2025)
Unleashing Hierarchical Reasoning: An LLM-Driven Framework for Training-Free Referring Video Object Segmentation
von: Zhao, Bingrui, et al.
Veröffentlicht: (2025)
von: Zhao, Bingrui, et al.
Veröffentlicht: (2025)
IFAdapter: Instance Feature Control for Grounded Text-to-Image Generation
von: Wu, Yinwei, et al.
Veröffentlicht: (2024)
von: Wu, Yinwei, et al.
Veröffentlicht: (2024)
Object-Attribute Binding in Text-to-Image Generation: Evaluation and Control
von: Trusca, Maria Mihaela, et al.
Veröffentlicht: (2024)
von: Trusca, Maria Mihaela, et al.
Veröffentlicht: (2024)
YOLO-Count: Differentiable Object Counting for Text-to-Image Generation
von: Zeng, Guanning, et al.
Veröffentlicht: (2025)
von: Zeng, Guanning, et al.
Veröffentlicht: (2025)
ComposeAnything: Composite Object Priors for Text-to-Image Generation
von: Khan, Zeeshan, et al.
Veröffentlicht: (2025)
von: Khan, Zeeshan, et al.
Veröffentlicht: (2025)
SpineCLUE: Automatic Vertebrae Identification Using Contrastive Learning and Uncertainty Estimation
von: Zhang, Sheng, et al.
Veröffentlicht: (2024)
von: Zhang, Sheng, et al.
Veröffentlicht: (2024)
Generating a Paracosm for Training-Free Zero-Shot Composed Image Retrieval
von: Wang, Tong, et al.
Veröffentlicht: (2026)
von: Wang, Tong, et al.
Veröffentlicht: (2026)
TelePhysics: Physics-Grounded Multi-Object Scene Generation from a Single Image with Real-Time Interaction
von: Zhang, Xin, et al.
Veröffentlicht: (2026)
von: Zhang, Xin, et al.
Veröffentlicht: (2026)
A Real-Time Framework for Domain-Adaptive Underwater Object Detection with Image Enhancement
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
Compass Control: Multi Object Orientation Control for Text-to-Image Generation
von: Parihar, Rishubh, et al.
Veröffentlicht: (2025)
von: Parihar, Rishubh, et al.
Veröffentlicht: (2025)
DeblurDiff: Real-World Image Deblurring with Generative Diffusion Models
von: Kong, Lingshun, et al.
Veröffentlicht: (2025)
von: Kong, Lingshun, et al.
Veröffentlicht: (2025)
ZoomLDM: Latent Diffusion Model for multi-scale image generation
von: Yellapragada, Srikar, et al.
Veröffentlicht: (2024)
von: Yellapragada, Srikar, et al.
Veröffentlicht: (2024)
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering
von: Zhang, Mingfang, et al.
Veröffentlicht: (2026)
von: Zhang, Mingfang, et al.
Veröffentlicht: (2026)
ST-SAM: SAM-Driven Self-Training Framework for Semi-Supervised Camouflaged Object Detection
von: Hu, Xihang, et al.
Veröffentlicht: (2025)
von: Hu, Xihang, et al.
Veröffentlicht: (2025)
Detecting and Classifying Defective Products in Images Using YOLO
von: Qi, Zhen, et al.
Veröffentlicht: (2024)
von: Qi, Zhen, et al.
Veröffentlicht: (2024)
A Physically-Grounded Attack and Adaptive Defense Framework for Real-World Low-Light Image Enhancement
von: Zhang, Tongshun, et al.
Veröffentlicht: (2026)
von: Zhang, Tongshun, et al.
Veröffentlicht: (2026)
CLIP-DFGS: A Hard Sample Mining Method for CLIP in Generalizable Person Re-Identification
von: Zhao, Huazhong, et al.
Veröffentlicht: (2024)
von: Zhao, Huazhong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Rethinking Referring Object Removal
von: Xue, Xiangtian, et al.
Veröffentlicht: (2024) -
Multiscale Low-Frequency Memory Network for Improved Feature Extraction in Convolutional Neural Networks
von: Wu, Fuzhi, et al.
Veröffentlicht: (2024) -
GGT-100K: Generative Ground Truth for Generalizable Real-World Image Restoration
von: Kong, Xiangtao, et al.
Veröffentlicht: (2026) -
TiFRe: Text-guided Video Frame Reduction for Efficient Video Multi-modal Large Language Models
von: Zheng, Xiangtian, et al.
Veröffentlicht: (2026) -
Intraoperative 2D/3D Registration via Spherical Similarity Learning and Differentiable Levenberg-Marquardt Optimization
von: Chen, Minheng, et al.
Veröffentlicht: (2025)