Gespeichert in:
| Hauptverfasser: | Wu, Jianzong, Hou, Liang, Yang, Haotian, Tao, Xin, Tian, Ye, Wan, Pengfei, Zhang, Di, Tong, Yunhai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2506.23858 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Does Hearing Help Seeing? Investigating Audio-Video Joint Denoising for Video Generation
von: Wu, Jianzong, et al.
Veröffentlicht: (2025)
von: Wu, Jianzong, et al.
Veröffentlicht: (2025)
Decouple and Track: Benchmarking and Improving Video Diffusion Transformers for Motion Transfer
von: Shi, Qingyu, et al.
Veröffentlicht: (2025)
von: Shi, Qingyu, et al.
Veröffentlicht: (2025)
DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation
von: Wu, Jianzong, et al.
Veröffentlicht: (2024)
von: Wu, Jianzong, et al.
Veröffentlicht: (2024)
Diffusion-Sharpening: Fine-tuning Diffusion Models with Denoising Trajectory Sharpening
von: Tian, Ye, et al.
Veröffentlicht: (2025)
von: Tian, Ye, et al.
Veröffentlicht: (2025)
VMonarch: Efficient Video Diffusion Transformers with Structured Attention
von: Liang, Cheng, et al.
Veröffentlicht: (2026)
von: Liang, Cheng, et al.
Veröffentlicht: (2026)
VideoTetris: Towards Compositional Text-to-Video Generation
von: Tian, Ye, et al.
Veröffentlicht: (2024)
von: Tian, Ye, et al.
Veröffentlicht: (2024)
MotionBooth: Motion-Aware Customized Text-to-Video Generation
von: Wu, Jianzong, et al.
Veröffentlicht: (2024)
von: Wu, Jianzong, et al.
Veröffentlicht: (2024)
Boosting Resolution Generalization of Diffusion Transformers with Randomized Positional Encodings
von: Hou, Liang, et al.
Veröffentlicht: (2025)
von: Hou, Liang, et al.
Veröffentlicht: (2025)
MMaDA: Multimodal Large Diffusion Language Models
von: Yang, Ling, et al.
Veröffentlicht: (2025)
von: Yang, Ling, et al.
Veröffentlicht: (2025)
MTV-Inpaint: Multi-Task Long Video Inpainting
von: Yang, Shiyuan, et al.
Veröffentlicht: (2025)
von: Yang, Shiyuan, et al.
Veröffentlicht: (2025)
Towards Language-Driven Video Inpainting via Multimodal Large Language Models
von: Wu, Jianzong, et al.
Veröffentlicht: (2024)
von: Wu, Jianzong, et al.
Veröffentlicht: (2024)
Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model
von: Shi, Qingyu, et al.
Veröffentlicht: (2025)
von: Shi, Qingyu, et al.
Veröffentlicht: (2025)
DreamRelation: Bridging Customization and Relation Generation
von: Shi, Qingyu, et al.
Veröffentlicht: (2024)
von: Shi, Qingyu, et al.
Veröffentlicht: (2024)
PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation
von: Liu, Zhuoman, et al.
Veröffentlicht: (2024)
von: Liu, Zhuoman, et al.
Veröffentlicht: (2024)
Decoupling Complexity from Scale in Latent Diffusion Model
von: Zhong, Tianxiong, et al.
Veröffentlicht: (2025)
von: Zhong, Tianxiong, et al.
Veröffentlicht: (2025)
Free Lunch Alignment of Text-to-Image Diffusion Models without Preference Image Pairs
von: Xian, Jia Jun Cheng, et al.
Veröffentlicht: (2025)
von: Xian, Jia Jun Cheng, et al.
Veröffentlicht: (2025)
One-Step Distillation of Discrete Diffusion Image Generators via Fixed-Point Iteration
von: Wang, Chaoyang, et al.
Veröffentlicht: (2026)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2026)
Training-free Diffusion Acceleration with Bottleneck Sampling
von: Tian, Ye, et al.
Veröffentlicht: (2025)
von: Tian, Ye, et al.
Veröffentlicht: (2025)
SALAD: Achieve High-Sparsity Attention via Efficient Linear Attention Tuning for Video Diffusion Transformer
von: Fang, Tongcheng, et al.
Veröffentlicht: (2026)
von: Fang, Tongcheng, et al.
Veröffentlicht: (2026)
VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization
von: Cheng, Junhao, et al.
Veröffentlicht: (2026)
von: Cheng, Junhao, et al.
Veröffentlicht: (2026)
MoBA: Mixture of Block Attention for Long-Context LLMs
von: Lu, Enzhe, et al.
Veröffentlicht: (2025)
von: Lu, Enzhe, et al.
Veröffentlicht: (2025)
FullDiT: Multi-Task Video Generative Foundation Model with Full Attention
von: Ju, Xuan, et al.
Veröffentlicht: (2025)
von: Ju, Xuan, et al.
Veröffentlicht: (2025)
BAGNet: A Boundary-Aware Graph Attention Network for 3D Point Cloud Semantic Segmentation
von: Tao, Wei, et al.
Veröffentlicht: (2025)
von: Tao, Wei, et al.
Veröffentlicht: (2025)
Triage: Hierarchical Visual Budgeting for Efficient Video Reasoning in Vision-Language Models
von: Wang, Anmin, et al.
Veröffentlicht: (2026)
von: Wang, Anmin, et al.
Veröffentlicht: (2026)
Direct-a-Video: Customized Video Generation with User-Directed Camera Movement and Object Motion
von: Yang, Shiyuan, et al.
Veröffentlicht: (2024)
von: Yang, Shiyuan, et al.
Veröffentlicht: (2024)
MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware Experts
von: Tao, Wei, et al.
Veröffentlicht: (2025)
von: Tao, Wei, et al.
Veröffentlicht: (2025)
Imbalance in Balance: Online Concept Balancing in Generation Models
von: Shi, Yukai, et al.
Veröffentlicht: (2025)
von: Shi, Yukai, et al.
Veröffentlicht: (2025)
I2V-Adapter: A General Image-to-Video Adapter for Diffusion Models
von: Guo, Xun, et al.
Veröffentlicht: (2023)
von: Guo, Xun, et al.
Veröffentlicht: (2023)
Motion Inversion for Video Customization
von: Wang, Luozhou, et al.
Veröffentlicht: (2024)
von: Wang, Luozhou, et al.
Veröffentlicht: (2024)
Efficient Video Diffusion Models: Advancements and Challenges
von: Shao, Shitong, et al.
Veröffentlicht: (2026)
von: Shao, Shitong, et al.
Veröffentlicht: (2026)
Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content
von: Wang, Qiuheng, et al.
Veröffentlicht: (2024)
von: Wang, Qiuheng, et al.
Veröffentlicht: (2024)
FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers
von: He, Xuanhua, et al.
Veröffentlicht: (2025)
von: He, Xuanhua, et al.
Veröffentlicht: (2025)
Stable Velocity: A Variance Perspective on Flow Matching
von: Yang, Donglin, et al.
Veröffentlicht: (2026)
von: Yang, Donglin, et al.
Veröffentlicht: (2026)
MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation
von: Tian, Ye, et al.
Veröffentlicht: (2025)
von: Tian, Ye, et al.
Veröffentlicht: (2025)
Towards Open Vocabulary Learning: A Survey
von: Wu, Jianzong, et al.
Veröffentlicht: (2023)
von: Wu, Jianzong, et al.
Veröffentlicht: (2023)
Owl-1: Omni World Model for Consistent Long Video Generation
von: Huang, Yuanhui, et al.
Veröffentlicht: (2024)
von: Huang, Yuanhui, et al.
Veröffentlicht: (2024)
BadVideo: Stealthy Backdoor Attack against Text-to-Video Generation
von: Wang, Ruotong, et al.
Veröffentlicht: (2025)
von: Wang, Ruotong, et al.
Veröffentlicht: (2025)
MemFlow: Flowing Adaptive Memory for Consistent and Efficient Long Video Narratives
von: Ji, Sihui, et al.
Veröffentlicht: (2025)
von: Ji, Sihui, et al.
Veröffentlicht: (2025)
Guiding Visual Autoregressive Models through Spectrum Weakening
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
PhysMaster: Mastering Physical Representation for Video Generation via Reinforcement Learning
von: Ji, Sihui, et al.
Veröffentlicht: (2025)
von: Ji, Sihui, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Does Hearing Help Seeing? Investigating Audio-Video Joint Denoising for Video Generation
von: Wu, Jianzong, et al.
Veröffentlicht: (2025) -
Decouple and Track: Benchmarking and Improving Video Diffusion Transformers for Motion Transfer
von: Shi, Qingyu, et al.
Veröffentlicht: (2025) -
DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation
von: Wu, Jianzong, et al.
Veröffentlicht: (2024) -
Diffusion-Sharpening: Fine-tuning Diffusion Models with Denoising Trajectory Sharpening
von: Tian, Ye, et al.
Veröffentlicht: (2025) -
VMonarch: Efficient Video Diffusion Transformers with Structured Attention
von: Liang, Cheng, et al.
Veröffentlicht: (2026)