Multi-focal Conditioned Latent Diffusion for Person Image Synthesis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Jiaqi, Zhang, Jichao, Rota, Paolo, Sebe, Nicu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
High-Fidelity 3D Facial Avatar Synthesis with Controllable Fine-Grained Expressions
von: He, Yikang, et al.
Veröffentlicht: (2026)
von: He, Yikang, et al.
Veröffentlicht: (2026)
Enhanced Multi-Scale Cross-Attention for Person Image Generation
von: Tang, Hao, et al.
Veröffentlicht: (2025)
von: Tang, Hao, et al.
Veröffentlicht: (2025)
Training and Tuning Generative Neural Radiance Fields for Attribute-Conditional 3D-Aware Face Generation
von: Zhang, Jichao, et al.
Veröffentlicht: (2022)
von: Zhang, Jichao, et al.
Veröffentlicht: (2022)
UVMap-ID: A Controllable and Personalized UV Map Generative Model
von: Wang, Weijie, et al.
Veröffentlicht: (2024)
von: Wang, Weijie, et al.
Veröffentlicht: (2024)
Reverse Personalization
von: Kung, Han-Wei, et al.
Veröffentlicht: (2025)
von: Kung, Han-Wei, et al.
Veröffentlicht: (2025)
TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation
von: Shu, Yan, et al.
Veröffentlicht: (2026)
von: Shu, Yan, et al.
Veröffentlicht: (2026)
Asymmetric GANs for Image-to-Image Translation
von: Tang, Hao, et al.
Veröffentlicht: (2019)
von: Tang, Hao, et al.
Veröffentlicht: (2019)
ESPLoRA: Enhanced Spatial Precision with Low-Rank Adaption in Text-to-Image Diffusion Models for High-Definition Synthesis
von: Rigo, Andrea, et al.
Veröffentlicht: (2025)
von: Rigo, Andrea, et al.
Veröffentlicht: (2025)
Anti-Forgetting Adaptation for Unsupervised Person Re-identification
von: Chen, Hao, et al.
Veröffentlicht: (2024)
von: Chen, Hao, et al.
Veröffentlicht: (2024)
Hallucination Early Detection in Diffusion Models
von: Betti, Federico, et al.
Veröffentlicht: (2026)
von: Betti, Federico, et al.
Veröffentlicht: (2026)
Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image Synthesis
von: Lu, Yanzuo, et al.
Veröffentlicht: (2024)
von: Lu, Yanzuo, et al.
Veröffentlicht: (2024)
CLIP is Strong Enough to Fight Back: Test-time Counterattacks towards Zero-shot Adversarial Robustness of CLIP
von: Xing, Songlong, et al.
Veröffentlicht: (2025)
von: Xing, Songlong, et al.
Veröffentlicht: (2025)
Spatial-Temporal Graph Mamba for Music-Guided Dance Video Synthesis
von: Tang, Hao, et al.
Veröffentlicht: (2025)
von: Tang, Hao, et al.
Veröffentlicht: (2025)
EarthMind: Leveraging Cross-Sensor Data for Advanced Earth Observation Interpretation with a Unified Multimodal LLM
von: Shu, Yan, et al.
Veröffentlicht: (2025)
von: Shu, Yan, et al.
Veröffentlicht: (2025)
Causal Disentanglement for Robust Long-tail Medical Image Generation
von: Nie, Weizhi, et al.
Veröffentlicht: (2025)
von: Nie, Weizhi, et al.
Veröffentlicht: (2025)
Optimizing Resource Consumption in Diffusion Models through Hallucination Early Detection
von: Betti, Federico, et al.
Veröffentlicht: (2024)
von: Betti, Federico, et al.
Veröffentlicht: (2024)
Inverse Virtual Try-On: Generating Multi-Category Product-Style Images from Clothed Individuals
von: Lobba, Davide, et al.
Veröffentlicht: (2025)
von: Lobba, Davide, et al.
Veröffentlicht: (2025)
Jointly Conditioned Diffusion Model for Multi-View Pose-Guided Person Image Synthesis
von: Xie, Chengyu, et al.
Veröffentlicht: (2025)
von: Xie, Chengyu, et al.
Veröffentlicht: (2025)
A Closer Look at Conditional Prompt Tuning for Vision-Language Models
von: Zhang, Ji, et al.
Veröffentlicht: (2025)
von: Zhang, Ji, et al.
Veröffentlicht: (2025)
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models
von: Li, Jinlong, et al.
Veröffentlicht: (2026)
von: Li, Jinlong, et al.
Veröffentlicht: (2026)
Stable-Hair v2: Real-World Hair Transfer via Multiple-View Diffusion Model
von: Sun, Kuiyuan, et al.
Veröffentlicht: (2025)
von: Sun, Kuiyuan, et al.
Veröffentlicht: (2025)
Generalized Fine-Grained Category Discovery with Multi-Granularity Conceptual Experts
von: Zheng, Haiyang, et al.
Veröffentlicht: (2025)
von: Zheng, Haiyang, et al.
Veröffentlicht: (2025)
RankFeat&RankWeight: Rank-1 Feature/Weight Removal for Out-of-distribution Detection
von: Song, Yue, et al.
Veröffentlicht: (2023)
von: Song, Yue, et al.
Veröffentlicht: (2023)
Cross-Modal and Uncertainty-Aware Agglomeration for Open-Vocabulary 3D Scene Understanding
von: Li, Jinlong, et al.
Veröffentlicht: (2025)
von: Li, Jinlong, et al.
Veröffentlicht: (2025)
Vision+X: A Survey on Multimodal Learning in the Light of Data
von: Zhu, Ye, et al.
Veröffentlicht: (2022)
von: Zhu, Ye, et al.
Veröffentlicht: (2022)
Hierarchical Cross-Attention Network for Virtual Try-On
von: Tang, Hao, et al.
Veröffentlicht: (2024)
von: Tang, Hao, et al.
Veröffentlicht: (2024)
Rethinking the Learning Paradigm for Facial Expression Recognition
von: Wang, Weijie, et al.
Veröffentlicht: (2022)
von: Wang, Weijie, et al.
Veröffentlicht: (2022)
Large Language Models for Multimodal Deformable Image Registration
von: Ma, Mingrui, et al.
Veröffentlicht: (2024)
von: Ma, Mingrui, et al.
Veröffentlicht: (2024)
AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual Grounding
von: Wang, Yidan, et al.
Veröffentlicht: (2025)
von: Wang, Yidan, et al.
Veröffentlicht: (2025)
CE-SDWV: Effective and Efficient Concept Erasure for Text-to-Image Diffusion Models via a Semantic-Driven Word Vocabulary
von: Tu, Jiahang, et al.
Veröffentlicht: (2025)
von: Tu, Jiahang, et al.
Veröffentlicht: (2025)
Diffusion-4K: Ultra-High-Resolution Image Synthesis with Latent Diffusion Models
von: Zhang, Jinjin, et al.
Veröffentlicht: (2025)
von: Zhang, Jinjin, et al.
Veröffentlicht: (2025)
NullFace: Training-Free Localized Face Anonymization
von: Kung, Han-Wei, et al.
Veröffentlicht: (2025)
von: Kung, Han-Wei, et al.
Veröffentlicht: (2025)
Graph Transformer GANs with Graph Masked Modeling for Architectural Layout Generation
von: Tang, Hao, et al.
Veröffentlicht: (2024)
von: Tang, Hao, et al.
Veröffentlicht: (2024)
Wasserstein-Aligned Hyperbolic Multi-View Clustering
von: Wang, Rui, et al.
Veröffentlicht: (2025)
von: Wang, Rui, et al.
Veröffentlicht: (2025)
Bilateral Reference for High-Resolution Dichotomous Image Segmentation
von: Zheng, Peng, et al.
Veröffentlicht: (2024)
von: Zheng, Peng, et al.
Veröffentlicht: (2024)
GradBias: Unveiling Word Influence on Bias in Text-to-Image Generative Models
von: D'Incà, Moreno, et al.
Veröffentlicht: (2024)
von: D'Incà, Moreno, et al.
Veröffentlicht: (2024)
Are Conditional Latent Diffusion Models Effective for Image Restoration?
von: Yuan, Yunchen, et al.
Veröffentlicht: (2024)
von: Yuan, Yunchen, et al.
Veröffentlicht: (2024)
Cues3D: Unleashing the Power of Sole NeRF for Consistent and Unique Instances in Open-Vocabulary 3D Panoptic Segmentation
von: Xue, Feng, et al.
Veröffentlicht: (2025)
von: Xue, Feng, et al.
Veröffentlicht: (2025)
RAGME: Retrieval Augmented Video Generation for Enhanced Motion Realism
von: Peruzzo, Elia, et al.
Veröffentlicht: (2025)
von: Peruzzo, Elia, et al.
Veröffentlicht: (2025)
Prototypical Hash Encoding for On-the-Fly Fine-Grained Category Discovery
von: Zheng, Haiyang, et al.
Veröffentlicht: (2024)
von: Zheng, Haiyang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
High-Fidelity 3D Facial Avatar Synthesis with Controllable Fine-Grained Expressions
von: He, Yikang, et al.
Veröffentlicht: (2026) -
Enhanced Multi-Scale Cross-Attention for Person Image Generation
von: Tang, Hao, et al.
Veröffentlicht: (2025) -
Training and Tuning Generative Neural Radiance Fields for Attribute-Conditional 3D-Aware Face Generation
von: Zhang, Jichao, et al.
Veröffentlicht: (2022) -
UVMap-ID: A Controllable and Personalized UV Map Generative Model
von: Wang, Weijie, et al.
Veröffentlicht: (2024) -
Reverse Personalization
von: Kung, Han-Wei, et al.
Veröffentlicht: (2025)