EZIGen: Enhancing zero-shot personalized image generation with precise subject encoding and decoupled guidance
Fuente:
arXiv
Saved in:
| Main Authors: | Duan, Zicheng, Ding, Yuxuan, Gou, Chenhui, Zhou, Ziqin, Smith, Ethan, Liu, Lingqiao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Training-Free Motion-Guided Video Generation with Enhanced Temporal Consistency Using Motion Consistency Loss
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
LiveWorld: Simulating Out-of-Sight Dynamics in Generative Video World Models
by: Duan, Zicheng, et al.
Published: (2026)
by: Duan, Zicheng, et al.
Published: (2026)
The CLIP Model is Secretly an Image-to-Prompt Converter
by: Ding, Yuxuan, et al.
Published: (2023)
by: Ding, Yuxuan, et al.
Published: (2023)
Points-to-3D: Structure-Aware 3D Generation with Point Cloud Priors
by: Xia, Jiatong, et al.
Published: (2026)
by: Xia, Jiatong, et al.
Published: (2026)
Mask-guided cross-image attention for zero-shot in-silico histopathologic image generation with a diffusion model
by: Winter, Dominik, et al.
Published: (2024)
by: Winter, Dominik, et al.
Published: (2024)
DFU: scale-robust diffusion model for zero-shot super-resolution image generation
by: Havrilla, Alex, et al.
Published: (2023)
by: Havrilla, Alex, et al.
Published: (2023)
Let Your Video Listen to Your Music!
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
Taming generative video models for zero-shot optical flow extraction
by: Kim, Seungwoo, et al.
Published: (2025)
by: Kim, Seungwoo, et al.
Published: (2025)
An Empirical Study on How Video-LLMs Answer Video Questions
by: Gou, Chenhui, et al.
Published: (2025)
by: Gou, Chenhui, et al.
Published: (2025)
Live image-based neurosurgical guidance and roadmap generation using unsupervised embedding
by: Sarwin, Gary, et al.
Published: (2023)
by: Sarwin, Gary, et al.
Published: (2023)
Retrieval-enriched zero-shot image classification in low-resource domains
by: Dall'Asen, Nicola, et al.
Published: (2024)
by: Dall'Asen, Nicola, et al.
Published: (2024)
Significantly improving zero-shot X-ray pathology classification via fine-tuning pre-trained image-text encoders
by: Jang, Jongseong, et al.
Published: (2022)
by: Jang, Jongseong, et al.
Published: (2022)
Source-Free Unsupervised Domain Adaptation with Hypothesis Consolidation of Prediction Rationale
by: Shu, Yangyang, et al.
Published: (2024)
by: Shu, Yangyang, et al.
Published: (2024)
Enhancing zero-shot learning in medical imaging: integrating clip with advanced techniques for improved chest x-ray analysis
by: Bhardwaj, Prakhar, et al.
Published: (2025)
by: Bhardwaj, Prakhar, et al.
Published: (2025)
Close-up-GS: Enhancing Close-Up View Synthesis in 3D Gaussian Splatting with Progressive Self-Training
by: Xia, Jiatong, et al.
Published: (2025)
by: Xia, Jiatong, et al.
Published: (2025)
HiTVideo: Hierarchical Tokenizers for Enhancing Text-to-Video Generation with Autoregressive Large Language Models
by: Zhou, Ziqin, et al.
Published: (2025)
by: Zhou, Ziqin, et al.
Published: (2025)
I2V3D: Controllable image-to-video generation with 3D guidance
by: Zhang, Zhiyuan, et al.
Published: (2025)
by: Zhang, Zhiyuan, et al.
Published: (2025)
IMAGE-ALCHEMY: Advancing subject fidelity in personalised text-to-image generation
by: Tiwari, Amritanshu, et al.
Published: (2025)
by: Tiwari, Amritanshu, et al.
Published: (2025)
Video models are zero-shot learners and reasoners
by: Wiedemer, Thaddäus, et al.
Published: (2025)
by: Wiedemer, Thaddäus, et al.
Published: (2025)
Efficient scene text image super-resolution with semantic guidance
by: TomyEnrique, LeoWu, et al.
Published: (2024)
by: TomyEnrique, LeoWu, et al.
Published: (2024)
Training-Free Instance-Aware 3D Scene Reconstruction and Diffusion-Based View Synthesis from Sparse Images
by: Xia, Jiatong, et al.
Published: (2026)
by: Xia, Jiatong, et al.
Published: (2026)
Enhancing Close-up Novel View Synthesis via Pseudo-labeling
by: Xia, Jiatong, et al.
Published: (2025)
by: Xia, Jiatong, et al.
Published: (2025)
TriSAM: Tri-Plane SAM for zero-shot cortical blood vessel segmentation in VEM images
by: Wan, Jia, et al.
Published: (2024)
by: Wan, Jia, et al.
Published: (2024)
VCP-CLIP: A visual context prompting model for zero-shot anomaly segmentation
by: Qu, Zhen, et al.
Published: (2024)
by: Qu, Zhen, et al.
Published: (2024)
Zero-shot Text-guided Infinite Image Synthesis with LLM guidance
by: Kwon, Soyeong, et al.
Published: (2024)
by: Kwon, Soyeong, et al.
Published: (2024)
Cross-Modal Attention Alignment Network with Auxiliary Text Description for zero-shot sketch-based image retrieval
by: Su, Hanwen, et al.
Published: (2024)
by: Su, Hanwen, et al.
Published: (2024)
A 1Mb mixed-precision quantized encoder for image classification and patch-based compression
by: Nguyen, Van Thien, et al.
Published: (2025)
by: Nguyen, Van Thien, et al.
Published: (2025)
SenCLIP: Enhancing zero-shot land-use mapping for Sentinel-2 with ground-level prompting
by: Jain, Pallavi, et al.
Published: (2024)
by: Jain, Pallavi, et al.
Published: (2024)
InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows
by: Ataallah, Kirolos, et al.
Published: (2024)
by: Ataallah, Kirolos, et al.
Published: (2024)
ESP-Zero: Unsupervised enhancement of zero-shot classification for Extremely Sparse Point cloud
by: Han, Jiayi, et al.
Published: (2024)
by: Han, Jiayi, et al.
Published: (2024)
MAEDAY: MAE for few and zero shot AnomalY-Detection
by: Schwartz, Eli, et al.
Published: (2022)
by: Schwartz, Eli, et al.
Published: (2022)
FADE: Few-shot/zero-shot Anomaly Detection Engine using Large Vision-Language Model
by: Li, Yuanwei, et al.
Published: (2024)
by: Li, Yuanwei, et al.
Published: (2024)
VQ-VA World: Towards High-Quality Visual Question-Visual Answering
by: Gou, Chenhui, et al.
Published: (2025)
by: Gou, Chenhui, et al.
Published: (2025)
Animalbooth: multimodal feature enhancement for animal subject personalization
by: Liu, Chen, et al.
Published: (2025)
by: Liu, Chen, et al.
Published: (2025)
Enhancing Fine-Grained Visual Recognition in the Low-Data Regime Through Feature Magnitude Regularization
by: Chapman, Avraham, et al.
Published: (2024)
by: Chapman, Avraham, et al.
Published: (2024)
FS-Diff: Semantic guidance and clarity-aware simultaneous multimodal image fusion and super-resolution
by: Jie, Yuchan, et al.
Published: (2025)
by: Jie, Yuchan, et al.
Published: (2025)
Adaptive few-shot learning for robust part quality classification in two-photon lithography
by: Jia, Sixian, et al.
Published: (2026)
by: Jia, Sixian, et al.
Published: (2026)
Compact single-shot ranging and near-far imaging using metasurfaces
by: Luo, Junjie, et al.
Published: (2026)
by: Luo, Junjie, et al.
Published: (2026)
MS-Diffusion: Multi-subject Zero-shot Image Personalization with Layout Guidance
by: Wang, Xierui, et al.
Published: (2024)
by: Wang, Xierui, et al.
Published: (2024)
Asynchronous Large Language Model Enhanced Planner for Autonomous Driving
by: Chen, Yuan, et al.
Published: (2024)
by: Chen, Yuan, et al.
Published: (2024)
Similar Items
-
Training-Free Motion-Guided Video Generation with Enhanced Temporal Consistency Using Motion Consistency Loss
by: Zhang, Xinyu, et al.
Published: (2025) -
LiveWorld: Simulating Out-of-Sight Dynamics in Generative Video World Models
by: Duan, Zicheng, et al.
Published: (2026) -
The CLIP Model is Secretly an Image-to-Prompt Converter
by: Ding, Yuxuan, et al.
Published: (2023) -
Points-to-3D: Structure-Aware 3D Generation with Point Cloud Priors
by: Xia, Jiatong, et al.
Published: (2026) -
Mask-guided cross-image attention for zero-shot in-silico histopathologic image generation with a diffusion model
by: Winter, Dominik, et al.
Published: (2024)