SkyReels-A2: Compose Anything in Video Diffusion Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fei, Zhengcong, Li, Debang, Qiu, Di, Wang, Jiahua, Dou, Yikun, Wang, Rui, Xu, Jingtao, Fan, Mingyuan, Chen, Guibin, Li, Yang, Zhou, Yahui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers
von: Fei, Zhengcong, et al.
Veröffentlicht: (2025)
von: Fei, Zhengcong, et al.
Veröffentlicht: (2025)
SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers
von: Qiu, Di, et al.
Veröffentlicht: (2025)
von: Qiu, Di, et al.
Veröffentlicht: (2025)
SkyReels-V3 Technique Report
von: Li, Debang, et al.
Veröffentlicht: (2026)
von: Li, Debang, et al.
Veröffentlicht: (2026)
SkyReels-V2: Infinite-length Film Generative Model
von: Chen, Guibin, et al.
Veröffentlicht: (2025)
von: Chen, Guibin, et al.
Veröffentlicht: (2025)
Video Diffusion Transformers are In-Context Learners
von: Fei, Zhengcong, et al.
Veröffentlicht: (2024)
von: Fei, Zhengcong, et al.
Veröffentlicht: (2024)
Ingredients: Blending Custom Photos with Video Diffusion Transformers
von: Fei, Zhengcong, et al.
Veröffentlicht: (2025)
von: Fei, Zhengcong, et al.
Veröffentlicht: (2025)
SkyReels-V4: Multi-modal Video-Audio Generation, Inpainting and Editing model
von: Chen, Guibin, et al.
Veröffentlicht: (2026)
von: Chen, Guibin, et al.
Veröffentlicht: (2026)
Scaling Diffusion Transformers to 16 Billion Parameters
von: Fei, Zhengcong, et al.
Veröffentlicht: (2024)
von: Fei, Zhengcong, et al.
Veröffentlicht: (2024)
Dimba: Transformer-Mamba Diffusion Models
von: Fei, Zhengcong, et al.
Veröffentlicht: (2024)
von: Fei, Zhengcong, et al.
Veröffentlicht: (2024)
Diffusion-RWKV: Scaling RWKV-Like Architectures for Diffusion Models
von: Fei, Zhengcong, et al.
Veröffentlicht: (2024)
von: Fei, Zhengcong, et al.
Veröffentlicht: (2024)
SkyReels-Text: Fine-Grained Font-Controllable Text Editing for Poster Design
von: Yu, Yunjie, et al.
Veröffentlicht: (2025)
von: Yu, Yunjie, et al.
Veröffentlicht: (2025)
Scalable Diffusion Models with State Space Backbone
von: Fei, Zhengcong, et al.
Veröffentlicht: (2024)
von: Fei, Zhengcong, et al.
Veröffentlicht: (2024)
A-JEPA: Joint-Embedding Predictive Architecture Can Listen
von: Fei, Zhengcong, et al.
Veröffentlicht: (2023)
von: Fei, Zhengcong, et al.
Veröffentlicht: (2023)
Music Consistency Models
von: Fei, Zhengcong, et al.
Veröffentlicht: (2024)
von: Fei, Zhengcong, et al.
Veröffentlicht: (2024)
FLUX that Plays Music
von: Fei, Zhengcong, et al.
Veröffentlicht: (2024)
von: Fei, Zhengcong, et al.
Veröffentlicht: (2024)
PodReels: Human-AI Co-Creation of Video Podcast Teasers
von: Wang, Sitong, et al.
Veröffentlicht: (2023)
von: Wang, Sitong, et al.
Veröffentlicht: (2023)
MovieCharacter: A Tuning-Free Framework for Controllable Character Video Synthesis
von: Qiu, Di, et al.
Veröffentlicht: (2024)
von: Qiu, Di, et al.
Veröffentlicht: (2024)
Personalize Anything for Free with Diffusion Transformer
von: Feng, Haoran, et al.
Veröffentlicht: (2025)
von: Feng, Haoran, et al.
Veröffentlicht: (2025)
ReelFramer: Human-AI Co-Creation for News-to-Video Translation
von: Wang, Sitong, et al.
Veröffentlicht: (2023)
von: Wang, Sitong, et al.
Veröffentlicht: (2023)
A Sanity Check on Composed Image Retrieval
von: Liu, Yikun, et al.
Veröffentlicht: (2026)
von: Liu, Yikun, et al.
Veröffentlicht: (2026)
CubeComposer: Spatio-Temporal Autoregressive 4K 360° Video Generation from Perspective Video
von: Li, Lingen, et al.
Veröffentlicht: (2026)
von: Li, Lingen, et al.
Veröffentlicht: (2026)
Systematic Evaluation and Guidelines for Segment Anything Model in Surgical Video Analysis
von: Yuan, Cheng, et al.
Veröffentlicht: (2024)
von: Yuan, Cheng, et al.
Veröffentlicht: (2024)
Zero-shot Composed Text-Image Retrieval
von: Liu, Yikun, et al.
Veröffentlicht: (2023)
von: Liu, Yikun, et al.
Veröffentlicht: (2023)
SayAnything: Audio-Driven Lip Synchronization with Conditional Video Diffusion
von: Ma, Junxian, et al.
Veröffentlicht: (2025)
von: Ma, Junxian, et al.
Veröffentlicht: (2025)
Anything in Any Scene: Photorealistic Video Object Insertion
von: Bai, Chen, et al.
Veröffentlicht: (2024)
von: Bai, Chen, et al.
Veröffentlicht: (2024)
Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
von: Huang, Xun, et al.
Veröffentlicht: (2025)
von: Huang, Xun, et al.
Veröffentlicht: (2025)
Get In Video: Add Anything You Want to the Video
von: Zhuang, Shaobin, et al.
Veröffentlicht: (2025)
von: Zhuang, Shaobin, et al.
Veröffentlicht: (2025)
AnimateAnything: Consistent and Controllable Animation for Video Generation
von: Lei, Guojun, et al.
Veröffentlicht: (2024)
von: Lei, Guojun, et al.
Veröffentlicht: (2024)
Tora: Trajectory-oriented Diffusion Transformer for Video Generation
von: Zhang, Zhenghao, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenghao, et al.
Veröffentlicht: (2024)
SAMEdge: An Edge-cloud Video Analytics Architecture for the Segment Anything Model
von: Lu, Rui, et al.
Veröffentlicht: (2024)
von: Lu, Rui, et al.
Veröffentlicht: (2024)
Growth Inhibitors for Suppressing Inappropriate Image Concepts in Diffusion Models
von: Chen, Die, et al.
Veröffentlicht: (2024)
von: Chen, Die, et al.
Veröffentlicht: (2024)
MoCapAnything: Unified 3D Motion Capture for Arbitrary Skeletons from Monocular Videos
von: Gong, Kehong, et al.
Veröffentlicht: (2025)
von: Gong, Kehong, et al.
Veröffentlicht: (2025)
ComposeAnything: Composite Object Priors for Text-to-Image Generation
von: Khan, Zeeshan, et al.
Veröffentlicht: (2025)
von: Khan, Zeeshan, et al.
Veröffentlicht: (2025)
DragAnything: Motion Control for Anything using Entity Representation
von: Wu, Weijia, et al.
Veröffentlicht: (2024)
von: Wu, Weijia, et al.
Veröffentlicht: (2024)
Responsible Diffusion Models via Constraining Text Embeddings within Safe Regions
von: Li, Zhiwen, et al.
Veröffentlicht: (2025)
von: Li, Zhiwen, et al.
Veröffentlicht: (2025)
Few-Step Diffusion via Score identity Distillation
von: Zhou, Mingyuan, et al.
Veröffentlicht: (2025)
von: Zhou, Mingyuan, et al.
Veröffentlicht: (2025)
Beyond Doping: Rare‐Earth Mediated Strategies for Rational Design of Multifunctional Electrocatalysts
von: Di Wang, et al.
Veröffentlicht: (2026)
von: Di Wang, et al.
Veröffentlicht: (2026)
Composable Effect Handling for Programming LLM-integrated Scripts
von: Wang, Di
Veröffentlicht: (2025)
von: Wang, Di
Veröffentlicht: (2025)
DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers
von: Wang, Lizhen, et al.
Veröffentlicht: (2025)
von: Wang, Lizhen, et al.
Veröffentlicht: (2025)
The Reel Deal: Designing and Evaluating LLM-Generated Short-Form Educational Videos
von: Stavrinou, Lazaros, et al.
Veröffentlicht: (2025)
von: Stavrinou, Lazaros, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers
von: Fei, Zhengcong, et al.
Veröffentlicht: (2025) -
SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers
von: Qiu, Di, et al.
Veröffentlicht: (2025) -
SkyReels-V3 Technique Report
von: Li, Debang, et al.
Veröffentlicht: (2026) -
SkyReels-V2: Infinite-length Film Generative Model
von: Chen, Guibin, et al.
Veröffentlicht: (2025) -
Video Diffusion Transformers are In-Context Learners
von: Fei, Zhengcong, et al.
Veröffentlicht: (2024)