MoCA: Identity-Preserving Text-to-Video Generation via Mixture of Cross Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xie, Qi, Ma, Yongjia, Di, Donglin, Gao, Xuehao, Yang, Xun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MoCA: Mixture-of-Components Attention for Scalable Compositional 3D Generation
von: Li, Zhiqi, et al.
Veröffentlicht: (2025)
von: Li, Zhiqi, et al.
Veröffentlicht: (2025)
Tuning-Free Long Video Generation via Global-Local Collaborative Diffusion
von: Ma, Yongjia, et al.
Veröffentlicht: (2025)
von: Ma, Yongjia, et al.
Veröffentlicht: (2025)
MoCA-Video: Motion-Aware Concept Alignment for Consistent Video Editing
von: Zhang, Tong, et al.
Veröffentlicht: (2025)
von: Zhang, Tong, et al.
Veröffentlicht: (2025)
TrAME: Trajectory-Anchored Multi-View Editing for Text-Guided 3D Gaussian Splatting Manipulation
von: Luo, Chaofan, et al.
Veröffentlicht: (2024)
von: Luo, Chaofan, et al.
Veröffentlicht: (2024)
UniCP: A Unified Caching and Pruning Framework for Efficient Video Generation
von: Sun, Wenzhang, et al.
Veröffentlicht: (2025)
von: Sun, Wenzhang, et al.
Veröffentlicht: (2025)
QR-LoRA: Efficient and Disentangled Fine-tuning via QR Decomposition for Customized Generation
von: Yang, Jiahui, et al.
Veröffentlicht: (2025)
von: Yang, Jiahui, et al.
Veröffentlicht: (2025)
EverybodyDance: Bipartite Graph-Based Identity Correspondence for Multi-Character Animation
von: Ling, Haotian, et al.
Veröffentlicht: (2025)
von: Ling, Haotian, et al.
Veröffentlicht: (2025)
DH-FaceVid-1K: A Large-Scale High-Quality Dataset for Face Video Generation
von: Di, Donglin, et al.
Veröffentlicht: (2024)
von: Di, Donglin, et al.
Veröffentlicht: (2024)
TV-3DG: Mastering Text-to-3D Customized Generation with Visual Prompt
von: Yang, Jiahui, et al.
Veröffentlicht: (2024)
von: Yang, Jiahui, et al.
Veröffentlicht: (2024)
MoCA3D: Monocular 3D Bounding Box Prediction in the Image Plane
von: Jeon, Changwoo, et al.
Veröffentlicht: (2026)
von: Jeon, Changwoo, et al.
Veröffentlicht: (2026)
Real Face Video Animation Platform
von: Chen, Xiaokai, et al.
Veröffentlicht: (2024)
von: Chen, Xiaokai, et al.
Veröffentlicht: (2024)
Hyper-3DG: Text-to-3D Gaussian Generation via Hypergraph
von: Di, Donglin, et al.
Veröffentlicht: (2024)
von: Di, Donglin, et al.
Veröffentlicht: (2024)
CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent Planning
von: Feng, He, et al.
Veröffentlicht: (2026)
von: Feng, He, et al.
Veröffentlicht: (2026)
MoSA: Motion-Coherent Human Video Generation via Structure-Appearance Decoupling
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement
von: Gao, Jiayi, et al.
Veröffentlicht: (2025)
von: Gao, Jiayi, et al.
Veröffentlicht: (2025)
Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation
von: Chen, Yuheng, et al.
Veröffentlicht: (2026)
von: Chen, Yuheng, et al.
Veröffentlicht: (2026)
One-Shot Pose-Driving Face Animation Platform
von: Feng, He, et al.
Veröffentlicht: (2024)
von: Feng, He, et al.
Veröffentlicht: (2024)
GUESS:GradUally Enriching SyntheSis for Text-Driven Human Motion Generation
von: Gao, Xuehao, et al.
Veröffentlicht: (2024)
von: Gao, Xuehao, et al.
Veröffentlicht: (2024)
Identity-Preserving Text-to-Video Generation by Frequency Decomposition
von: Yuan, Shenghai, et al.
Veröffentlicht: (2024)
von: Yuan, Shenghai, et al.
Veröffentlicht: (2024)
CA-IDD: Cross-Attention Guided Identity-Conditional Diffusion for Identity-Consistent Face Swapping
von: Rana, Md Shohel, et al.
Veröffentlicht: (2026)
von: Rana, Md Shohel, et al.
Veröffentlicht: (2026)
MoEE: Mixture of Emotion Experts for Audio-Driven Portrait Animation
von: Liu, Huaize, et al.
Veröffentlicht: (2025)
von: Liu, Huaize, et al.
Veröffentlicht: (2025)
MoGA: Mixture-of-Groups Attention for End-to-End Long Video Generation
von: Jia, Weinan, et al.
Veröffentlicht: (2025)
von: Jia, Weinan, et al.
Veröffentlicht: (2025)
Adams Bashforth Moulton Solver for Inversion and Editing in Rectified Flow
von: Ma, Yongjia, et al.
Veröffentlicht: (2025)
von: Ma, Yongjia, et al.
Veröffentlicht: (2025)
DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation
von: Feng, He, et al.
Veröffentlicht: (2025)
von: Feng, He, et al.
Veröffentlicht: (2025)
FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation
von: Wang, Yuanzhi, et al.
Veröffentlicht: (2026)
von: Wang, Yuanzhi, et al.
Veröffentlicht: (2026)
Identity-Preserving Text-to-Video Generation Guided by Simple yet Effective Spatial-Temporal Decoupled Representations
von: Wang, Yuji, et al.
Veröffentlicht: (2025)
von: Wang, Yuji, et al.
Veröffentlicht: (2025)
WildActor: Unconstrained Identity-Preserving Video Generation
von: Guo, Qin, et al.
Veröffentlicht: (2026)
von: Guo, Qin, et al.
Veröffentlicht: (2026)
Identity-Preserving Image-to-Video Generation via Reward-Guided Optimization
von: Shen, Liao, et al.
Veröffentlicht: (2025)
von: Shen, Liao, et al.
Veröffentlicht: (2025)
CA^2ST: Cross-Attention in Audio, Space, and Time for Holistic Video Recognition
von: Lee, Jongseo, et al.
Veröffentlicht: (2025)
von: Lee, Jongseo, et al.
Veröffentlicht: (2025)
Knowledge Priors for Identity-Disentangled Open-Set Privacy-Preserving Video FER
von: Xu, Feng, et al.
Veröffentlicht: (2026)
von: Xu, Feng, et al.
Veröffentlicht: (2026)
Multimodal Cancer Survival Analysis via Hypergraph Learning with Cross-Modality Rebalance
von: Qu, Mingcheng, et al.
Veröffentlicht: (2025)
von: Qu, Mingcheng, et al.
Veröffentlicht: (2025)
BIG-MoE: Bypass Isolated Gating MoE for Generalized Multimodal Face Anti-Spoofing
von: Ma, Yingjie, et al.
Veröffentlicht: (2024)
von: Ma, Yingjie, et al.
Veröffentlicht: (2024)
DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation
von: Chen, Junhao, et al.
Veröffentlicht: (2025)
von: Chen, Junhao, et al.
Veröffentlicht: (2025)
MoGAN: Improving Motion Quality in Video Diffusion via Few-Step Motion Adversarial Post-Training
von: Xue, Haotian, et al.
Veröffentlicht: (2025)
von: Xue, Haotian, et al.
Veröffentlicht: (2025)
VMoBA: Mixture-of-Block Attention for Video Diffusion Models
von: Wu, Jianzong, et al.
Veröffentlicht: (2025)
von: Wu, Jianzong, et al.
Veröffentlicht: (2025)
ConsisDrive: Identity-Preserving Driving World Models for Video Generation by Instance Mask
von: Yang, Zhuoran, et al.
Veröffentlicht: (2026)
von: Yang, Zhuoran, et al.
Veröffentlicht: (2026)
EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion
von: Wei, Jiangchuan, et al.
Veröffentlicht: (2025)
von: Wei, Jiangchuan, et al.
Veröffentlicht: (2025)
EigenActor: Variant Body-Object Interaction Generation Evolved from Invariant Action Basis Reasoning
von: Gao, Xuehao, et al.
Veröffentlicht: (2025)
von: Gao, Xuehao, et al.
Veröffentlicht: (2025)
Memory-Augmented Incomplete Multimodal Survival Prediction via Cross-Slide and Gene-Attentive Hypergraph Learning
von: Qu, Mingcheng, et al.
Veröffentlicht: (2025)
von: Qu, Mingcheng, et al.
Veröffentlicht: (2025)
Text2QR: Harmonizing Aesthetic Customization and Scanning Robustness for Text-Guided QR Code Generation
von: Wu, Guangyang, et al.
Veröffentlicht: (2024)
von: Wu, Guangyang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MoCA: Mixture-of-Components Attention for Scalable Compositional 3D Generation
von: Li, Zhiqi, et al.
Veröffentlicht: (2025) -
Tuning-Free Long Video Generation via Global-Local Collaborative Diffusion
von: Ma, Yongjia, et al.
Veröffentlicht: (2025) -
MoCA-Video: Motion-Aware Concept Alignment for Consistent Video Editing
von: Zhang, Tong, et al.
Veröffentlicht: (2025) -
TrAME: Trajectory-Anchored Multi-View Editing for Text-Guided 3D Gaussian Splatting Manipulation
von: Luo, Chaofan, et al.
Veröffentlicht: (2024) -
UniCP: A Unified Caching and Pruning Framework for Efficient Video Generation
von: Sun, Wenzhang, et al.
Veröffentlicht: (2025)