Multivariate Diffusion Transformer with Decoupled Attention for High-Fidelity Mask-Text Collaborative Facial Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Cao, Yushe, Shi, Dianxi, Fu, Xing, Zou, Xuechao, Peng, Haikuo, Li, Xueqi, Yu, Chun, Xing, Junliang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mixture of Global and Local Experts with Diffusion Transformer for Controllable Face Generation
by: Zou, Xuechao, et al.
Published: (2025)
by: Zou, Xuechao, et al.
Published: (2025)
High-Fidelity Lake Extraction via Two-Stage Prompt Enhancement: Establishing a Novel Baseline and Benchmark
by: Chen, Ben, et al.
Published: (2023)
by: Chen, Ben, et al.
Published: (2023)
A Parallel Attention Network for Cattle Face Recognition
by: Li, Jiayu, et al.
Published: (2024)
by: Li, Jiayu, et al.
Published: (2024)
UV-Mamba: A DCN-Enhanced State Space Model for Urban Village Boundary Identification in High-Resolution Remote Sensing Images
by: Li, Lulin, et al.
Published: (2024)
by: Li, Lulin, et al.
Published: (2024)
LEFormer: A Hybrid CNN-Transformer Architecture for Accurate Lake Extraction from Remote Sensing Imagery
by: Chen, Ben, et al.
Published: (2023)
by: Chen, Ben, et al.
Published: (2023)
ARFA: An Asymmetric Receptive Field Autoencoder Model for Spatiotemporal Prediction
by: Zhang, Wenxuan, et al.
Published: (2023)
by: Zhang, Wenxuan, et al.
Published: (2023)
Hierarchical Fusion of Local and Global Visual Features with Mixture-of-Experts for Remote Sensing Image Scene Classification
by: Tang, Yuanhao, et al.
Published: (2025)
by: Tang, Yuanhao, et al.
Published: (2025)
DeCoDe: Defer-and-Complement Decision-Making via Decoupled Concept Bottleneck Models
by: He, Chengbo, et al.
Published: (2025)
by: He, Chengbo, et al.
Published: (2025)
MDiffFR: Modality-Guided Diffusion Generation for Cold-start Items in Federated Recommendation
by: Fu, Kang, et al.
Published: (2025)
by: Fu, Kang, et al.
Published: (2025)
CEIDM: A Controlled Entity and Interaction Diffusion Model for Enhanced Text-to-Image Generation
by: Yang, Mingyue, et al.
Published: (2025)
by: Yang, Mingyue, et al.
Published: (2025)
Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion
by: Li, Lijiang, et al.
Published: (2026)
by: Li, Lijiang, et al.
Published: (2026)
Adapting Vision Foundation Models for Robust Cloud Segmentation in Remote Sensing Images
by: Zou, Xuechao, et al.
Published: (2024)
by: Zou, Xuechao, et al.
Published: (2024)
Dynamic Dictionary Learning for Remote Sensing Image Segmentation
by: Zou, Xuechao, et al.
Published: (2025)
by: Zou, Xuechao, et al.
Published: (2025)
Enhancing LLM Reasoning with Multi-Path Collaborative Reactive and Reflection agents
by: He, Chengbo, et al.
Published: (2024)
by: He, Chengbo, et al.
Published: (2024)
UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation
by: Kang, Wonjun, et al.
Published: (2025)
by: Kang, Wonjun, et al.
Published: (2025)
Masked Diffusion Generative Recommendation
by: Mu, Lingyu, et al.
Published: (2026)
by: Mu, Lingyu, et al.
Published: (2026)
CA-Edit: Causality-Aware Condition Adapter for High-Fidelity Local Facial Attribute Editing
by: Xian, Xiaole, et al.
Published: (2024)
by: Xian, Xiaole, et al.
Published: (2024)
High-Fidelity Relightable Monocular Portrait Animation with Lighting-Controllable Video Diffusion Model
by: Guo, Mingtao, et al.
Published: (2025)
by: Guo, Mingtao, et al.
Published: (2025)
Face-MakeUp: Multimodal Facial Prompts for Text-to-Image Generation
by: Dai, Dawei, et al.
Published: (2025)
by: Dai, Dawei, et al.
Published: (2025)
FundusGAN: A Hierarchical Feature-Aware Generative Framework for High-Fidelity Fundus Image Generation
by: Hou, Qingshan, et al.
Published: (2025)
by: Hou, Qingshan, et al.
Published: (2025)
High-Fidelity Diffusion Face Swapping with ID-Constrained Facial Conditioning
by: He, Dailan, et al.
Published: (2025)
by: He, Dailan, et al.
Published: (2025)
Nonlinear Optical Second Harmonic Generation Characteristics in Cylindrical GaAs/Ga1–ηAlηAs Quantum Dots
by: Xiaolong Yan, et al.
Published: (2024)
by: Xiaolong Yan, et al.
Published: (2024)
Domain Adaptive Attention Learning for Unsupervised Person Re-Identification
by: Huang, Yangru, et al.
Published: (2019)
by: Huang, Yangru, et al.
Published: (2019)
MagDiff: Multi-Alignment Diffusion for High-Fidelity Video Generation and Editing
by: Zhao, Haoyu, et al.
Published: (2023)
by: Zhao, Haoyu, et al.
Published: (2023)
GPTFace: Generative Pre-training of Facial-Linguistic Transformer by Span Masking and Weakly Correlated Text-image Data
by: Li, Yudong, et al.
Published: (2025)
by: Li, Yudong, et al.
Published: (2025)
Text-Conditioned Diffusion Model for High-Fidelity Korean Font Generation
by: Sami, Abdul, et al.
Published: (2025)
by: Sami, Abdul, et al.
Published: (2025)
Navigating Large-Pose Challenge for High-Fidelity Face Reenactment with Video Diffusion Model
by: Guo, Mingtao, et al.
Published: (2025)
by: Guo, Mingtao, et al.
Published: (2025)
PACT: Proactive Asking for Continual Task Assistance in Human-Robot Collaboration
by: He, Chengbo, et al.
Published: (2026)
by: He, Chengbo, et al.
Published: (2026)
An FPGA-Based Reconfigurable Accelerator for Convolution-Transformer Hybrid EfficientViT
by: Shao, Haikuo, et al.
Published: (2024)
by: Shao, Haikuo, et al.
Published: (2024)
AudioGAN: A Compact and Efficient Framework for Real-Time High-Fidelity Text-to-Audio Generation
by: Chung, HaeChun
Published: (2025)
by: Chung, HaeChun
Published: (2025)
Trio-ViT: Post-Training Quantization and Acceleration for Softmax-Free Efficient Vision Transformer
by: Shi, Huihong, et al.
Published: (2024)
by: Shi, Huihong, et al.
Published: (2024)
Unifying and Enhancing Graph Transformers via a Hierarchical Mask Framework
by: Xing, Yujie, et al.
Published: (2025)
by: Xing, Yujie, et al.
Published: (2025)
MARPO: A Reflective Policy Optimization for Multi Agent Reinforcement Learning
by: Wu, Cuiling, et al.
Published: (2025)
by: Wu, Cuiling, et al.
Published: (2025)
LaMamba-Diff: Linear-Time High-Fidelity Diffusion Models Based on Local Attention and Mamba
by: Fu, Yunxiang, et al.
Published: (2024)
by: Fu, Yunxiang, et al.
Published: (2024)
From Text to Mask: Localizing Entities Using the Attention of Text-to-Image Diffusion Models
by: Xiao, Changming, et al.
Published: (2023)
by: Xiao, Changming, et al.
Published: (2023)
FreEformer: Frequency Enhanced Transformer for Multivariate Time Series Forecasting
by: Yue, Wenzhen, et al.
Published: (2025)
by: Yue, Wenzhen, et al.
Published: (2025)
Emotic Masked Autoencoder with Attention Fusion for Facial Expression Recognition
by: Nguyen-Xuan, Bach, et al.
Published: (2024)
by: Nguyen-Xuan, Bach, et al.
Published: (2024)
Multivariate Fidelities
by: Nuradha, Theshani, et al.
Published: (2024)
by: Nuradha, Theshani, et al.
Published: (2024)
High-Fidelity and Long-Duration Human Image Animation with Diffusion Transformer
by: Zheng, Shen, et al.
Published: (2025)
by: Zheng, Shen, et al.
Published: (2025)
FD2Talk: Towards Generalized Talking Head Generation with Facial Decoupled Diffusion Model
by: Yao, Ziyu, et al.
Published: (2024)
by: Yao, Ziyu, et al.
Published: (2024)
Similar Items
-
Mixture of Global and Local Experts with Diffusion Transformer for Controllable Face Generation
by: Zou, Xuechao, et al.
Published: (2025) -
High-Fidelity Lake Extraction via Two-Stage Prompt Enhancement: Establishing a Novel Baseline and Benchmark
by: Chen, Ben, et al.
Published: (2023) -
A Parallel Attention Network for Cattle Face Recognition
by: Li, Jiayu, et al.
Published: (2024) -
UV-Mamba: A DCN-Enhanced State Space Model for Urban Village Boundary Identification in High-Resolution Remote Sensing Images
by: Li, Lulin, et al.
Published: (2024) -
LEFormer: A Hybrid CNN-Transformer Architecture for Accurate Lake Extraction from Remote Sensing Imagery
by: Chen, Ben, et al.
Published: (2023)