Enhanced Multi-Scale Cross-Attention for Person Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Hao, Shao, Ling, Sebe, Nicu, Van Gool, Luc |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Graph Transformer GANs with Graph Masked Modeling for Architectural Layout Generation
by: Tang, Hao, et al.
Published: (2024)
by: Tang, Hao, et al.
Published: (2024)
Spatial-Temporal Graph Mamba for Music-Guided Dance Video Synthesis
by: Tang, Hao, et al.
Published: (2025)
by: Tang, Hao, et al.
Published: (2025)
Hierarchical Cross-Attention Network for Virtual Try-On
by: Tang, Hao, et al.
Published: (2024)
by: Tang, Hao, et al.
Published: (2024)
Asymmetric GANs for Image-to-Image Translation
by: Tang, Hao, et al.
Published: (2019)
by: Tang, Hao, et al.
Published: (2019)
Multi-focal Conditioned Latent Diffusion for Person Image Synthesis
by: Liu, Jiaqi, et al.
Published: (2025)
by: Liu, Jiaqi, et al.
Published: (2025)
CityLoc: 6DoF Pose Distributional Localization for Text Descriptions in Large-Scale Scenes with Gaussian Representation
by: Ma, Qi, et al.
Published: (2025)
by: Ma, Qi, et al.
Published: (2025)
EarthMind: Leveraging Cross-Sensor Data for Advanced Earth Observation Interpretation with a Unified Multimodal LLM
by: Shu, Yan, et al.
Published: (2025)
by: Shu, Yan, et al.
Published: (2025)
Reverse Personalization
by: Kung, Han-Wei, et al.
Published: (2025)
by: Kung, Han-Wei, et al.
Published: (2025)
Sharing Key Semantics in Transformer Makes Efficient Image Restoration
by: Ren, Bin, et al.
Published: (2024)
by: Ren, Bin, et al.
Published: (2024)
Investigating the Effectiveness of Cross-Attention to Unlock Zero-Shot Editing of Text-to-Video Diffusion Models
by: Motamed, Saman, et al.
Published: (2024)
by: Motamed, Saman, et al.
Published: (2024)
Fully-Geometric Cross-Attention for Point Cloud Registration
by: Wang, Weijie, et al.
Published: (2025)
by: Wang, Weijie, et al.
Published: (2025)
Efficient Degradation-agnostic Image Restoration via Channel-Wise Functional Decomposition and Manifold Regularization
by: Ren, Bin, et al.
Published: (2025)
by: Ren, Bin, et al.
Published: (2025)
Anti-Forgetting Adaptation for Unsupervised Person Re-identification
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
Retrieval Augmented Generation and Understanding in Vision: A Survey and New Outlook
by: Zheng, Xu, et al.
Published: (2025)
by: Zheng, Xu, et al.
Published: (2025)
ShapeSplat: A Large-scale Dataset of Gaussian Splats and Their Self-Supervised Pretraining
by: Ma, Qi, et al.
Published: (2024)
by: Ma, Qi, et al.
Published: (2024)
HandDiff: 3D Hand Pose Estimation with Diffusion on Image-Point Cloud
by: Cheng, Wencan, et al.
Published: (2024)
by: Cheng, Wencan, et al.
Published: (2024)
PVAFN: Point-Voxel Attention Fusion Network with Multi-Pooling Enhancing for 3D Object Detection
by: Li, Yidi, et al.
Published: (2024)
by: Li, Yidi, et al.
Published: (2024)
Key-Graph Transformer for Image Restoration
by: Ren, Bin, et al.
Published: (2024)
by: Ren, Bin, et al.
Published: (2024)
Any Image Restoration via Efficient Spatial-Frequency Degradation Adaptation
by: Ren, Bin, et al.
Published: (2025)
by: Ren, Bin, et al.
Published: (2025)
Textual Knowledge Matters: Cross-Modality Co-Teaching for Generalized Visual Class Discovery
by: Zheng, Haiyang, et al.
Published: (2024)
by: Zheng, Haiyang, et al.
Published: (2024)
Inverse Virtual Try-On: Generating Multi-Category Product-Style Images from Clothed Individuals
by: Lobba, Davide, et al.
Published: (2025)
by: Lobba, Davide, et al.
Published: (2025)
Bringing Masked Autoencoders Explicit Contrastive Properties for Point Cloud Self-Supervised Learning
by: Ren, Bin, et al.
Published: (2024)
by: Ren, Bin, et al.
Published: (2024)
RAGME: Retrieval Augmented Video Generation for Enhanced Motion Realism
by: Peruzzo, Elia, et al.
Published: (2025)
by: Peruzzo, Elia, et al.
Published: (2025)
Cross-Modal and Uncertainty-Aware Agglomeration for Open-Vocabulary 3D Scene Understanding
by: Li, Jinlong, et al.
Published: (2025)
by: Li, Jinlong, et al.
Published: (2025)
Towards Online Real-Time Memory-based Video Inpainting Transformers
by: Thiry, Guillaume, et al.
Published: (2024)
by: Thiry, Guillaume, et al.
Published: (2024)
Test-time Training for Hyperspectral Image Super-resolution
by: Li, Ke, et al.
Published: (2024)
by: Li, Ke, et al.
Published: (2024)
Generalized Fine-Grained Category Discovery with Multi-Granularity Conceptual Experts
by: Zheng, Haiyang, et al.
Published: (2025)
by: Zheng, Haiyang, et al.
Published: (2025)
Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encoding
by: Li, Yue, et al.
Published: (2025)
by: Li, Yue, et al.
Published: (2025)
Lego: Learning to Disentangle and Invert Personalized Concepts Beyond Object Appearance in Text-to-Image Diffusion Models
by: Motamed, Saman, et al.
Published: (2023)
by: Motamed, Saman, et al.
Published: (2023)
Implicit-Zoo: A Large-Scale Dataset of Neural Implicit Functions for 2D Images and 3D Scenes
by: Ma, Qi, et al.
Published: (2024)
by: Ma, Qi, et al.
Published: (2024)
Training and Tuning Generative Neural Radiance Fields for Attribute-Conditional 3D-Aware Face Generation
by: Zhang, Jichao, et al.
Published: (2022)
by: Zhang, Jichao, et al.
Published: (2022)
Transferable-guided Attention Is All You Need for Video Domain Adaptation
by: Sacilotti, André, et al.
Published: (2024)
by: Sacilotti, André, et al.
Published: (2024)
CLIP is Strong Enough to Fight Back: Test-time Counterattacks towards Zero-shot Adversarial Robustness of CLIP
by: Xing, Songlong, et al.
Published: (2025)
by: Xing, Songlong, et al.
Published: (2025)
Causal Disentanglement for Robust Long-tail Medical Image Generation
by: Nie, Weizhi, et al.
Published: (2025)
by: Nie, Weizhi, et al.
Published: (2025)
MCANet: Medical Image Segmentation with Multi-Scale Cross-Axis Attention
by: Shao, Hao, et al.
Published: (2023)
by: Shao, Hao, et al.
Published: (2023)
Reducing Unimodal Bias in Multi-Modal Semantic Segmentation with Multi-Scale Functional Entropy Regularization
by: Zheng, Xu, et al.
Published: (2025)
by: Zheng, Xu, et al.
Published: (2025)
UVMap-ID: A Controllable and Personalized UV Map Generative Model
by: Wang, Weijie, et al.
Published: (2024)
by: Wang, Weijie, et al.
Published: (2024)
MatIR: A Hybrid Mamba-Transformer Image Restoration Model
by: Wen, Juan, et al.
Published: (2025)
by: Wen, Juan, et al.
Published: (2025)
GradBias: Unveiling Word Influence on Bias in Text-to-Image Generative Models
by: D'Incà, Moreno, et al.
Published: (2024)
by: D'Incà, Moreno, et al.
Published: (2024)
ESPLoRA: Enhanced Spatial Precision with Low-Rank Adaption in Text-to-Image Diffusion Models for High-Definition Synthesis
by: Rigo, Andrea, et al.
Published: (2025)
by: Rigo, Andrea, et al.
Published: (2025)
Similar Items
-
Graph Transformer GANs with Graph Masked Modeling for Architectural Layout Generation
by: Tang, Hao, et al.
Published: (2024) -
Spatial-Temporal Graph Mamba for Music-Guided Dance Video Synthesis
by: Tang, Hao, et al.
Published: (2025) -
Hierarchical Cross-Attention Network for Virtual Try-On
by: Tang, Hao, et al.
Published: (2024) -
Asymmetric GANs for Image-to-Image Translation
by: Tang, Hao, et al.
Published: (2019) -
Multi-focal Conditioned Latent Diffusion for Person Image Synthesis
by: Liu, Jiaqi, et al.
Published: (2025)