Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Wenjing, Yang, Huan, Tuo, Zixi, He, Huiguo, Zhu, Junchen, Fu, Jianlong, Liu, Jiaying |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Zero-Reference Low-Light Enhancement via Physical Quadruple Priors
by: Wang, Wenjing, et al.
Published: (2024)
by: Wang, Wenjing, et al.
Published: (2024)
DreamStory: Open-Domain Story Visualization by LLM-Guided Multi-Subject Consistent Diffusion
by: He, Huiguo, et al.
Published: (2024)
by: He, Huiguo, et al.
Published: (2024)
Improving Multi-Subject Consistency in Open-Domain Image Generation with Isolation and Reposition Attention
by: He, Huiguo, et al.
Published: (2024)
by: He, Huiguo, et al.
Published: (2024)
Solving Diffusion ODEs with Optimal Boundary Conditions for Better Image Super-Resolution
by: Ma, Yiyang, et al.
Published: (2023)
by: Ma, Yiyang, et al.
Published: (2023)
DiffInk: Glyph- and Style-Aware Latent Diffusion Transformer for Text to Online Handwriting Generation
by: Pan, Wei, et al.
Published: (2025)
by: Pan, Wei, et al.
Published: (2025)
RoboSwap: A GAN-driven Video Diffusion Framework For Unsupervised Robot Arm Swapping
by: Bai, Yang, et al.
Published: (2025)
by: Bai, Yang, et al.
Published: (2025)
Learning Position-Aware Implicit Neural Network for Real-World Face Inpainting
by: Zhao, Bo, et al.
Published: (2024)
by: Zhao, Bo, et al.
Published: (2024)
SwapAnyone: Consistent and Realistic Video Synthesis for Swapping Any Person into Any Video
by: Zhao, Chengshu, et al.
Published: (2025)
by: Zhao, Chengshu, et al.
Published: (2025)
Understanding Attention Mechanism in Video Diffusion Models
by: Liu, Bingyan, et al.
Published: (2025)
by: Liu, Bingyan, et al.
Published: (2025)
HiFiVFS: High Fidelity Video Face Swapping
by: Chen, Xu, et al.
Published: (2024)
by: Chen, Xu, et al.
Published: (2024)
AICL: Action In-Context Learning for Video Diffusion Model
by: Liu, Jianzhi, et al.
Published: (2024)
by: Liu, Jianzhi, et al.
Published: (2024)
Dual-Stream Diffusion Net for Text-to-Video Generation
by: Liu, Binhui, et al.
Published: (2023)
by: Liu, Binhui, et al.
Published: (2023)
GaussianSwap: Animatable Video Face Swapping with 3D Gaussian Splatting
by: Cheng, Xuan, et al.
Published: (2026)
by: Cheng, Xuan, et al.
Published: (2026)
UniCtrl: Improving the Spatiotemporal Consistency of Text-to-Video Diffusion Models via Training-Free Unified Attention Control
by: Xia, Tian, et al.
Published: (2024)
by: Xia, Tian, et al.
Published: (2024)
Language-Guided Face Animation by Recurrent StyleGAN-based Generator
by: Hang, Tiankai, et al.
Published: (2022)
by: Hang, Tiankai, et al.
Published: (2022)
CanonSwap: High-Fidelity and Consistent Video Face Swapping via Canonical Space Modulation
by: Luo, Xiangyang, et al.
Published: (2025)
by: Luo, Xiangyang, et al.
Published: (2025)
Mod-Adapter: Tuning-Free and Versatile Multi-concept Personalization via Modulation Adapter
by: Zhong, Weizhi, et al.
Published: (2025)
by: Zhong, Weizhi, et al.
Published: (2025)
OpenDance: Multimodal Controllable 3D Dance Generation with Large-scale Internet Data
by: Zhang, Jinlu, et al.
Published: (2025)
by: Zhang, Jinlu, et al.
Published: (2025)
Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation
by: Yan, Xin, et al.
Published: (2024)
by: Yan, Xin, et al.
Published: (2024)
VRMDiff: Text-Guided Video Referring Matting Generation of Diffusion
by: Yang, Lehan, et al.
Published: (2025)
by: Yang, Lehan, et al.
Published: (2025)
ContextDrag: Precise Drag-Based Image Editing via Context-Preserving Token Injection and Position-Aligned Attention
by: He, Huiguo, et al.
Published: (2025)
by: He, Huiguo, et al.
Published: (2025)
Face Swap via Diffusion Model
by: Wang, Feifei
Published: (2024)
by: Wang, Feifei
Published: (2024)
DirectSwap: Mask-Free Cross-Identity Training and Benchmarking for Expression-Consistent Video Head Swapping
by: Wang, Yanan, et al.
Published: (2025)
by: Wang, Yanan, et al.
Published: (2025)
DreamSwapV: Mask-guided Subject Swapping for Any Customized Video Editing
by: Wang, Weitao, et al.
Published: (2025)
by: Wang, Weitao, et al.
Published: (2025)
Training-Free Semantic Video Composition via Pre-trained Diffusion Model
by: Guo, Jiaqi, et al.
Published: (2024)
by: Guo, Jiaqi, et al.
Published: (2024)
High-Fidelity Diffusion Face Swapping with ID-Constrained Facial Conditioning
by: He, Dailan, et al.
Published: (2025)
by: He, Dailan, et al.
Published: (2025)
StarGen: A Spatiotemporal Autoregression Framework with Video Diffusion Model for Scalable and Controllable Scene Generation
by: Zhai, Shangjin, et al.
Published: (2025)
by: Zhai, Shangjin, et al.
Published: (2025)
AID: Attention Interpolation of Text-to-Image Diffusion
by: He, Qiyuan, et al.
Published: (2024)
by: He, Qiyuan, et al.
Published: (2024)
AnyText: Multilingual Visual Text Generation And Editing
by: Tuo, Yuxiang, et al.
Published: (2023)
by: Tuo, Yuxiang, et al.
Published: (2023)
Time2General: Learning Spatiotemporal Invariant Representations for Domain-Generalization Video Semantic Segmentation
by: Chen, Siyu, et al.
Published: (2026)
by: Chen, Siyu, et al.
Published: (2026)
PTDiffusion: Free Lunch for Generating Optical Illusion Hidden Pictures with Phase-Transferred Diffusion Model
by: Gao, Xiang, et al.
Published: (2025)
by: Gao, Xiang, et al.
Published: (2025)
HOI-Swap: Swapping Objects in Videos with Hand-Object Interaction Awareness
by: Xue, Zihui, et al.
Published: (2024)
by: Xue, Zihui, et al.
Published: (2024)
VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation
by: Ren, Weiming, et al.
Published: (2024)
by: Ren, Weiming, et al.
Published: (2024)
VividFace: A Diffusion-Based Hybrid Framework for High-Fidelity Video Face Swapping
by: Shao, Hao, et al.
Published: (2024)
by: Shao, Hao, et al.
Published: (2024)
AlignVid: Training-Free Attention Scaling for Semantic Fidelity in Text-Guided Image-to-Video Generation
by: Liu, Yexin, et al.
Published: (2025)
by: Liu, Yexin, et al.
Published: (2025)
VFace: A Training-Free Approach for Diffusion-Based Video Face Swapping
by: Baliah, Sanoojan, et al.
Published: (2026)
by: Baliah, Sanoojan, et al.
Published: (2026)
CSTA: CNN-based Spatiotemporal Attention for Video Summarization
by: Son, Jaewon, et al.
Published: (2024)
by: Son, Jaewon, et al.
Published: (2024)
TP2O: Creative Text Pair-to-Object Generation using Balance Swap-Sampling
by: Li, Jun, et al.
Published: (2023)
by: Li, Jun, et al.
Published: (2023)
AnyText2: Visual Text Generation and Editing With Customizable Attributes
by: Tuo, Yuxiang, et al.
Published: (2024)
by: Tuo, Yuxiang, et al.
Published: (2024)
CustomVideoX: 3D Reference Attention Driven Dynamic Adaptation for Zero-Shot Customized Video Diffusion Transformers
by: She, D., et al.
Published: (2025)
by: She, D., et al.
Published: (2025)
Similar Items
-
Zero-Reference Low-Light Enhancement via Physical Quadruple Priors
by: Wang, Wenjing, et al.
Published: (2024) -
DreamStory: Open-Domain Story Visualization by LLM-Guided Multi-Subject Consistent Diffusion
by: He, Huiguo, et al.
Published: (2024) -
Improving Multi-Subject Consistency in Open-Domain Image Generation with Isolation and Reposition Attention
by: He, Huiguo, et al.
Published: (2024) -
Solving Diffusion ODEs with Optimal Boundary Conditions for Better Image Super-Resolution
by: Ma, Yiyang, et al.
Published: (2023) -
DiffInk: Glyph- and Style-Aware Latent Diffusion Transformer for Text to Online Handwriting Generation
by: Pan, Wei, et al.
Published: (2025)