Salvato in:
| Autori principali: | He, Jing, Li, Haodong, Hu, Yongzhe, Shen, Guibao, Cai, Yingjie, Qiu, Weichao, Chen, Ying-Cong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2410.02067 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
OmniBooth: Learning Latent Control for Image Synthesis with Multi-modal Instruction
di: Li, Leheng, et al.
Pubblicazione: (2024)
di: Li, Leheng, et al.
Pubblicazione: (2024)
SyntheOcc: Synthesize Geometric-Controlled Street View Images through 3D Semantic MPIs
di: Li, Leheng, et al.
Pubblicazione: (2024)
di: Li, Leheng, et al.
Pubblicazione: (2024)
Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model
di: He, Jing, et al.
Pubblicazione: (2025)
di: He, Jing, et al.
Pubblicazione: (2025)
StereoPilot: Learning Unified and Efficient Stereo Conversion via Generative Priors
di: Shen, Guibao, et al.
Pubblicazione: (2025)
di: Shen, Guibao, et al.
Pubblicazione: (2025)
PRM: Photometric Stereo based Large Reconstruction Model
di: Ge, Wenhang, et al.
Pubblicazione: (2024)
di: Ge, Wenhang, et al.
Pubblicazione: (2024)
Motion Inversion for Video Customization
di: Wang, Luozhou, et al.
Pubblicazione: (2024)
di: Wang, Luozhou, et al.
Pubblicazione: (2024)
Text-Anchored Score Composition: Tackling Condition Misalignment in Text-to-Image Diffusion Models
di: Wang, Luozhou, et al.
Pubblicazione: (2023)
di: Wang, Luozhou, et al.
Pubblicazione: (2023)
DisPose: Disentangling Pose Guidance for Controllable Human Image Animation
di: Li, Hongxiang, et al.
Pubblicazione: (2024)
di: Li, Hongxiang, et al.
Pubblicazione: (2024)
Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing
di: Zhao, Xiangyu, et al.
Pubblicazione: (2025)
di: Zhao, Xiangyu, et al.
Pubblicazione: (2025)
SG-Adapter: Enhancing Text-to-Image Generation with Scene Graph Guidance
di: Shen, Guibao, et al.
Pubblicazione: (2024)
di: Shen, Guibao, et al.
Pubblicazione: (2024)
IC-Custom: Diverse Image Customization via In-Context Learning
di: Li, Yaowei, et al.
Pubblicazione: (2025)
di: Li, Yaowei, et al.
Pubblicazione: (2025)
DisCo: Disentangled Control for Realistic Human Dance Generation
di: Wang, Tan, et al.
Pubblicazione: (2023)
di: Wang, Tan, et al.
Pubblicazione: (2023)
An Item is Worth a Prompt: Versatile Image Editing with Disentangled Control
di: Feng, Aosong, et al.
Pubblicazione: (2024)
di: Feng, Aosong, et al.
Pubblicazione: (2024)
Learning Disentangled Identifiers for Action-Customized Text-to-Image Generation
di: Huang, Siteng, et al.
Pubblicazione: (2023)
di: Huang, Siteng, et al.
Pubblicazione: (2023)
Prompt-Agnostic Adversarial Perturbation for Customized Diffusion Models
di: Wan, Cong, et al.
Pubblicazione: (2024)
di: Wan, Cong, et al.
Pubblicazione: (2024)
CamPilot: Improving Camera Control in Video Diffusion Model with Efficient Camera Reward Feedback
di: Ge, Wenhang, et al.
Pubblicazione: (2026)
di: Ge, Wenhang, et al.
Pubblicazione: (2026)
RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data
di: Chen, Harold Haodong, et al.
Pubblicazione: (2026)
di: Chen, Harold Haodong, et al.
Pubblicazione: (2026)
TAVP: Task-Adaptive Visual Prompt for Cross-domain Few-shot Segmentation
di: Yang, Jiaqi, et al.
Pubblicazione: (2024)
di: Yang, Jiaqi, et al.
Pubblicazione: (2024)
Visual Prompt-Agnostic Evolution
di: Wang, Junze, et al.
Pubblicazione: (2026)
di: Wang, Junze, et al.
Pubblicazione: (2026)
RectifiedHR: Enable Efficient High-Resolution Synthesis via Energy Rectification
di: Yang, Zhen, et al.
Pubblicazione: (2025)
di: Yang, Zhen, et al.
Pubblicazione: (2025)
Paragraph-to-Image Generation with Information-Enriched Diffusion Model
di: Wu, Weijia, et al.
Pubblicazione: (2023)
di: Wu, Weijia, et al.
Pubblicazione: (2023)
DisMo: Disentangled Motion Representations for Open-World Motion Transfer
di: Ressler-Antal, Thomas, et al.
Pubblicazione: (2025)
di: Ressler-Antal, Thomas, et al.
Pubblicazione: (2025)
DisFaceRep: Representation Disentanglement for Co-occurring Facial Components in Weakly Supervised Face Parsing
di: Wang, Xiaoqin, et al.
Pubblicazione: (2025)
di: Wang, Xiaoqin, et al.
Pubblicazione: (2025)
Lotus: Diffusion-based Visual Foundation Model for High-quality Dense Prediction
di: He, Jing, et al.
Pubblicazione: (2024)
di: He, Jing, et al.
Pubblicazione: (2024)
Attend and Enrich: Enhanced Visual Prompt for Zero-Shot Learning
di: Liu, Man, et al.
Pubblicazione: (2024)
di: Liu, Man, et al.
Pubblicazione: (2024)
DisControlFace: Adding Disentangled Control to Diffusion Autoencoder for One-shot Explicit Facial Image Editing
di: Jia, Haozhe, et al.
Pubblicazione: (2023)
di: Jia, Haozhe, et al.
Pubblicazione: (2023)
Bi-TTA: Bidirectional Test-Time Adapter for Remote Physiological Measurement
di: Li, Haodong, et al.
Pubblicazione: (2024)
di: Li, Haodong, et al.
Pubblicazione: (2024)
Generalizing Alignment Paradigm of Text-to-Image Generation with Preferences through $f$-divergence Minimization
di: Sun, Haoyuan, et al.
Pubblicazione: (2024)
di: Sun, Haoyuan, et al.
Pubblicazione: (2024)
Semantic-Enriched Latent Visual Reasoning
di: Xu, Tianrun, et al.
Pubblicazione: (2026)
di: Xu, Tianrun, et al.
Pubblicazione: (2026)
T-Rex-Omni: Integrating Negative Visual Prompt in Generic Object Detection
di: Zhou, Jiazhou, et al.
Pubblicazione: (2025)
di: Zhou, Jiazhou, et al.
Pubblicazione: (2025)
Go with Your Gut: Scaling Confidence for Autoregressive Image Generation
di: Chen, Harold Haodong, et al.
Pubblicazione: (2025)
di: Chen, Harold Haodong, et al.
Pubblicazione: (2025)
TV-3DG: Mastering Text-to-3D Customized Generation with Visual Prompt
di: Yang, Jiahui, et al.
Pubblicazione: (2024)
di: Yang, Jiahui, et al.
Pubblicazione: (2024)
FreeCustom: Tuning-Free Customized Image Generation for Multi-Concept Composition
di: Ding, Ganggui, et al.
Pubblicazione: (2024)
di: Ding, Ganggui, et al.
Pubblicazione: (2024)
OmniPrism: Learning Disentangled Visual Concept for Image Generation
di: Li, Yangyang, et al.
Pubblicazione: (2024)
di: Li, Yangyang, et al.
Pubblicazione: (2024)
IF-Bench: Benchmarking and Enhancing MLLMs for Infrared Images with Generative Visual Prompting
di: Zhang, Tao, et al.
Pubblicazione: (2025)
di: Zhang, Tao, et al.
Pubblicazione: (2025)
PET-DINO: Unifying Visual Cues into Grounding DINO with Prompt-Enriched Training
di: Fu, Weifu, et al.
Pubblicazione: (2026)
di: Fu, Weifu, et al.
Pubblicazione: (2026)
AttenCraft: Attention-guided Disentanglement of Multiple Concepts for Text-to-Image Customization
di: Shentu, Junjie, et al.
Pubblicazione: (2024)
di: Shentu, Junjie, et al.
Pubblicazione: (2024)
GroundingBooth: Grounding Text-to-Image Customization
di: Xiong, Zhexiao, et al.
Pubblicazione: (2024)
di: Xiong, Zhexiao, et al.
Pubblicazione: (2024)
SCC-YOLO: An Improved Object Detector for Assisting in Brain Tumor Diagnosis
di: Bai, Runci, et al.
Pubblicazione: (2025)
di: Bai, Runci, et al.
Pubblicazione: (2025)
Advancing high-fidelity 3D and Texture Generation with 2.5D latents
di: Yang, Xin, et al.
Pubblicazione: (2025)
di: Yang, Xin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
OmniBooth: Learning Latent Control for Image Synthesis with Multi-modal Instruction
di: Li, Leheng, et al.
Pubblicazione: (2024) -
SyntheOcc: Synthesize Geometric-Controlled Street View Images through 3D Semantic MPIs
di: Li, Leheng, et al.
Pubblicazione: (2024) -
Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model
di: He, Jing, et al.
Pubblicazione: (2025) -
StereoPilot: Learning Unified and Efficient Stereo Conversion via Generative Priors
di: Shen, Guibao, et al.
Pubblicazione: (2025) -
PRM: Photometric Stereo based Large Reconstruction Model
di: Ge, Wenhang, et al.
Pubblicazione: (2024)