GroundingBooth: Grounding Text-to-Image Customization
Fuente:
arXiv
Salvato in:
| Autori principali: | Xiong, Zhexiao, Xiong, Wei, Shi, Jing, Zhang, He, Song, Yizhi, Jacobs, Nathan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
PhysAlign: Physics-Coherent Image-to-Video Generation through Feature and 3D Representation Alignment
di: Xiong, Zhexiao, et al.
Pubblicazione: (2026)
di: Xiong, Zhexiao, et al.
Pubblicazione: (2026)
Towards Open-World Generation of Stereo Images and Unsupervised Matching
di: Qiao, Feng, et al.
Pubblicazione: (2025)
di: Qiao, Feng, et al.
Pubblicazione: (2025)
PanoDreamer: Consistent Text to 360-Degree Scene Generation
di: Xiong, Zhexiao, et al.
Pubblicazione: (2025)
di: Xiong, Zhexiao, et al.
Pubblicazione: (2025)
DeclutterNeRF: Generative-Free 3D Scene Recovery for Occlusion Removal
di: Liu, Wanzhou, et al.
Pubblicazione: (2025)
di: Liu, Wanzhou, et al.
Pubblicazione: (2025)
Mixed-View Panorama Synthesis using Geospatially Guided Diffusion
di: Xiong, Zhexiao, et al.
Pubblicazione: (2024)
di: Xiong, Zhexiao, et al.
Pubblicazione: (2024)
GenOpticalFlow: A Generative Approach to Unsupervised Optical Flow Learning
di: Luo, Yixuan, et al.
Pubblicazione: (2026)
di: Luo, Yixuan, et al.
Pubblicazione: (2026)
MotionBooth: Motion-Aware Customized Text-to-Video Generation
di: Wu, Jianzong, et al.
Pubblicazione: (2024)
di: Wu, Jianzong, et al.
Pubblicazione: (2024)
MultiBooth: Towards Generating All Your Concepts in an Image from Text
di: Zhu, Chenyang, et al.
Pubblicazione: (2024)
di: Zhu, Chenyang, et al.
Pubblicazione: (2024)
UniDrive-WM: Unified Understanding, Planning and Generation World Model For Autonomous Driving
di: Xiong, Zhexiao, et al.
Pubblicazione: (2026)
di: Xiong, Zhexiao, et al.
Pubblicazione: (2026)
AttnDreamBooth: Towards Text-Aligned Personalized Text-to-Image Generation
di: Pang, Lianyu, et al.
Pubblicazione: (2024)
di: Pang, Lianyu, et al.
Pubblicazione: (2024)
Dynamic Prompting of Frozen Text-to-Image Diffusion Models for Panoptic Narrative Grounding
di: Li, Hongyu, et al.
Pubblicazione: (2024)
di: Li, Hongyu, et al.
Pubblicazione: (2024)
Floating No More: Object-Ground Reconstruction from a Single Image
di: Man, Yunze, et al.
Pubblicazione: (2024)
di: Man, Yunze, et al.
Pubblicazione: (2024)
PersonaBooth: Personalized Text-to-Motion Generation
di: Kim, Boeun, et al.
Pubblicazione: (2025)
di: Kim, Boeun, et al.
Pubblicazione: (2025)
DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation
di: Chen, Hong, et al.
Pubblicazione: (2023)
di: Chen, Hong, et al.
Pubblicazione: (2023)
SceneBooth: Diffusion-based Framework for Subject-preserved Text-to-Image Generation
di: Chai, Shang, et al.
Pubblicazione: (2025)
di: Chai, Shang, et al.
Pubblicazione: (2025)
AutoFocus: Uncertainty-Aware Active Visual Search for GUI Grounding
di: Yao, Ruilin, et al.
Pubblicazione: (2026)
di: Yao, Ruilin, et al.
Pubblicazione: (2026)
PSM: Learning Probabilistic Embeddings for Multi-scale Zero-Shot Soundscape Mapping
di: Khanal, Subash, et al.
Pubblicazione: (2024)
di: Khanal, Subash, et al.
Pubblicazione: (2024)
InstructBooth: Instruction-following Personalized Text-to-Image Generation
di: Chae, Daewon, et al.
Pubblicazione: (2023)
di: Chae, Daewon, et al.
Pubblicazione: (2023)
OmniBooth: Learning Latent Control for Image Synthesis with Multi-modal Instruction
di: Li, Leheng, et al.
Pubblicazione: (2024)
di: Li, Leheng, et al.
Pubblicazione: (2024)
Grounding-IQA: Grounding Multimodal Language Model for Image Quality Assessment
di: Chen, Zheng, et al.
Pubblicazione: (2024)
di: Chen, Zheng, et al.
Pubblicazione: (2024)
StyleBooth: Image Style Editing with Multimodal Instruction
di: Han, Zhen, et al.
Pubblicazione: (2024)
di: Han, Zhen, et al.
Pubblicazione: (2024)
RealCustom: Narrowing Real Text Word for Real-Time Open-Domain Text-to-Image Customization
di: Huang, Mengqi, et al.
Pubblicazione: (2024)
di: Huang, Mengqi, et al.
Pubblicazione: (2024)
MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models
di: Hua, Hang, et al.
Pubblicazione: (2025)
di: Hua, Hang, et al.
Pubblicazione: (2025)
Compositional Image-Text Matching and Retrieval by Grounding Entities
di: Vongala, Madhukar Reddy, et al.
Pubblicazione: (2025)
di: Vongala, Madhukar Reddy, et al.
Pubblicazione: (2025)
Image Difference Grounding with Natural Language
di: Wang, Wenxuan, et al.
Pubblicazione: (2025)
di: Wang, Wenxuan, et al.
Pubblicazione: (2025)
UGround: Towards Unified Visual Grounding with Unrolled Transformers
di: Qian, Rui, et al.
Pubblicazione: (2025)
di: Qian, Rui, et al.
Pubblicazione: (2025)
HCMA: Hierarchical Cross-model Alignment for Grounded Text-to-Image Generation
di: Wang, Hang, et al.
Pubblicazione: (2025)
di: Wang, Hang, et al.
Pubblicazione: (2025)
ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning
di: Guo, Yuxiang, et al.
Pubblicazione: (2025)
di: Guo, Yuxiang, et al.
Pubblicazione: (2025)
Style Customization of Text-to-Vector Generation with Image Diffusion Priors
di: Zhang, Peiying, et al.
Pubblicazione: (2025)
di: Zhang, Peiying, et al.
Pubblicazione: (2025)
Visual Grounding with Multi-modal Conditional Adaptation
di: Yao, Ruilin, et al.
Pubblicazione: (2024)
di: Yao, Ruilin, et al.
Pubblicazione: (2024)
PhyCustom: Towards Realistic Physical Customization in Text-to-Image Generation
di: Wu, Fan, et al.
Pubblicazione: (2025)
di: Wu, Fan, et al.
Pubblicazione: (2025)
GVDIFF: Grounded Text-to-Video Generation with Diffusion Models
di: Dou, Huanzhang, et al.
Pubblicazione: (2024)
di: Dou, Huanzhang, et al.
Pubblicazione: (2024)
Open-Text Aerial Detection: A Unified Framework For Aerial Visual Grounding And Detection
di: Wei, Guoting, et al.
Pubblicazione: (2026)
di: Wei, Guoting, et al.
Pubblicazione: (2026)
Learning to Customize Text-to-Image Diffusion In Diverse Context
di: Kim, Taewook, et al.
Pubblicazione: (2024)
di: Kim, Taewook, et al.
Pubblicazione: (2024)
Q-Ground: Image Quality Grounding with Large Multi-modality Models
di: Chen, Chaofeng, et al.
Pubblicazione: (2024)
di: Chen, Chaofeng, et al.
Pubblicazione: (2024)
GIE-Bench: Towards Grounded Evaluation for Text-Guided Image Editing
di: Qian, Yusu, et al.
Pubblicazione: (2025)
di: Qian, Yusu, et al.
Pubblicazione: (2025)
HyperDreamBooth: HyperNetworks for Fast Personalization of Text-to-Image Models
di: Ruiz, Nataniel, et al.
Pubblicazione: (2023)
di: Ruiz, Nataniel, et al.
Pubblicazione: (2023)
Emergent Visual Grounding in Large Multimodal Models Without Grounding Supervision
di: Cao, Shengcao, et al.
Pubblicazione: (2024)
di: Cao, Shengcao, et al.
Pubblicazione: (2024)
From Objects to Anywhere: A Holistic Benchmark for Multi-level Visual Grounding in 3D Scenes
di: Wang, Tianxu, et al.
Pubblicazione: (2025)
di: Wang, Tianxu, et al.
Pubblicazione: (2025)
Content-Style Decoupling for Unsupervised Makeup Transfer without Generating Pseudo Ground Truth
di: Sun, Zhaoyang, et al.
Pubblicazione: (2024)
di: Sun, Zhaoyang, et al.
Pubblicazione: (2024)
Documenti analoghi
-
PhysAlign: Physics-Coherent Image-to-Video Generation through Feature and 3D Representation Alignment
di: Xiong, Zhexiao, et al.
Pubblicazione: (2026) -
Towards Open-World Generation of Stereo Images and Unsupervised Matching
di: Qiao, Feng, et al.
Pubblicazione: (2025) -
PanoDreamer: Consistent Text to 360-Degree Scene Generation
di: Xiong, Zhexiao, et al.
Pubblicazione: (2025) -
DeclutterNeRF: Generative-Free 3D Scene Recovery for Occlusion Removal
di: Liu, Wanzhou, et al.
Pubblicazione: (2025) -
Mixed-View Panorama Synthesis using Geospatially Guided Diffusion
di: Xiong, Zhexiao, et al.
Pubblicazione: (2024)