Faster Diffusion via Temporal Attention Decomposition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Haozhe, Zhang, Wentian, Xie, Jinheng, Faccio, Francesco, Xu, Mengmeng, Xiang, Tao, Shou, Mike Zheng, Perez-Rua, Juan-Manuel, Schmidhuber, Jürgen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning Useful Representations of Recurrent Neural Network Weight Matrices
von: Herrmann, Vincent, et al.
Veröffentlicht: (2024)
von: Herrmann, Vincent, et al.
Veröffentlicht: (2024)
Lazy Layers to Make Fine-Tuned Diffusion Models More Traceable
von: Liu, Haozhe, et al.
Veröffentlicht: (2024)
von: Liu, Haozhe, et al.
Veröffentlicht: (2024)
Show-o2: Improved Native Unified Multimodal Models
von: Xie, Jinheng, et al.
Veröffentlicht: (2025)
von: Xie, Jinheng, et al.
Veröffentlicht: (2025)
WMAdapter: Adding WaterMark Control to Latent Diffusion Models
von: Ci, Hai, et al.
Veröffentlicht: (2024)
von: Ci, Hai, et al.
Veröffentlicht: (2024)
MarDini: Masked Autoregressive Diffusion for Video Generation at Scale
von: Liu, Haozhe, et al.
Veröffentlicht: (2024)
von: Liu, Haozhe, et al.
Veröffentlicht: (2024)
TPDiff: Temporal Pyramid Video Diffusion Model
von: Ran, Lingmin, et al.
Veröffentlicht: (2025)
von: Ran, Lingmin, et al.
Veröffentlicht: (2025)
Highway Value Iteration Networks
von: Wang, Yuhui, et al.
Veröffentlicht: (2024)
von: Wang, Yuhui, et al.
Veröffentlicht: (2024)
Highway Reinforcement Learning
von: Wang, Yuhui, et al.
Veröffentlicht: (2024)
von: Wang, Yuhui, et al.
Veröffentlicht: (2024)
Hyper-VolTran: Fast and Generalizable One-Shot Image to 3D Object Structure via HyperNetworks
von: Simon, Christian, et al.
Veröffentlicht: (2023)
von: Simon, Christian, et al.
Veröffentlicht: (2023)
Dynamically Masked Discriminator for Generative Adversarial Networks
von: Zhang, Wentian, et al.
Veröffentlicht: (2023)
von: Zhang, Wentian, et al.
Veröffentlicht: (2023)
D-AR: Diffusion via Autoregressive Models
von: Gao, Ziteng, et al.
Veröffentlicht: (2025)
von: Gao, Ziteng, et al.
Veröffentlicht: (2025)
Language Agents as Optimizable Graphs
von: Zhuge, Mingchen, et al.
Veröffentlicht: (2024)
von: Zhuge, Mingchen, et al.
Veröffentlicht: (2024)
Upside Down Reinforcement Learning with Policy Generators
von: Di Ventura, Jacopo, et al.
Veröffentlicht: (2025)
von: Di Ventura, Jacopo, et al.
Veröffentlicht: (2025)
Learning Flow Fields in Attention for Controllable Person Image Generation
von: Zhou, Zijian, et al.
Veröffentlicht: (2024)
von: Zhou, Zijian, et al.
Veröffentlicht: (2024)
GenTron: Diffusion Transformers for Image and Video Generation
von: Chen, Shoufa, et al.
Veröffentlicht: (2023)
von: Chen, Shoufa, et al.
Veröffentlicht: (2023)
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing
von: Piękos, Piotr, et al.
Veröffentlicht: (2025)
von: Piękos, Piotr, et al.
Veröffentlicht: (2025)
Can Video Diffusion Model Reconstruct 4D Geometry?
von: Mai, Jinjie, et al.
Veröffentlicht: (2025)
von: Mai, Jinjie, et al.
Veröffentlicht: (2025)
On the Convergence and Stability of Upside-Down Reinforcement Learning, Goal-Conditioned Supervised Learning, and Online Decision Transformers
von: Štrupl, Miroslav, et al.
Veröffentlicht: (2025)
von: Štrupl, Miroslav, et al.
Veröffentlicht: (2025)
Curious Causality-Seeking Agents Learn Meta Causal World
von: Zhao, Zhiyu, et al.
Veröffentlicht: (2025)
von: Zhao, Zhiyu, et al.
Veröffentlicht: (2025)
LayerTracer: Cognitive-Aligned Layered SVG Synthesis via Diffusion Transformer
von: Song, Yiren, et al.
Veröffentlicht: (2025)
von: Song, Yiren, et al.
Veröffentlicht: (2025)
Scaling Value Iteration Networks to 5000 Layers for Extreme Long-Term Planning
von: Wang, Yuhui, et al.
Veröffentlicht: (2024)
von: Wang, Yuhui, et al.
Veröffentlicht: (2024)
Towards a Robust Soft Baby Robot With Rich Interaction Ability for Advanced Machine Learning Algorithms
von: Alhakami, Mohannad, et al.
Veröffentlicht: (2024)
von: Alhakami, Mohannad, et al.
Veröffentlicht: (2024)
How to Correctly do Semantic Backpropagation on Language-based Agentic Systems
von: Wang, Wenyi, et al.
Veröffentlicht: (2024)
von: Wang, Wenyi, et al.
Veröffentlicht: (2024)
Learning Long-form Video Prior via Generative Pre-Training
von: Xie, Jinheng, et al.
Veröffentlicht: (2024)
von: Xie, Jinheng, et al.
Veröffentlicht: (2024)
Adaptive Caching for Faster Video Generation with Diffusion Transformers
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing
von: Cong, Yuren, et al.
Veröffentlicht: (2023)
von: Cong, Yuren, et al.
Veröffentlicht: (2023)
DiffSim: Taming Diffusion Models for Evaluating Visual Similarity
von: Song, Yiren, et al.
Veröffentlicht: (2024)
von: Song, Yiren, et al.
Veröffentlicht: (2024)
Edit2Perceive: Image Editing Diffusion Models Are Strong Dense Perceivers
von: Shi, Yiqing, et al.
Veröffentlicht: (2025)
von: Shi, Yiqing, et al.
Veröffentlicht: (2025)
MakeAnything: Harnessing Diffusion Transformers for Multi-Domain Procedural Sequence Generation
von: Song, Yiren, et al.
Veröffentlicht: (2025)
von: Song, Yiren, et al.
Veröffentlicht: (2025)
La generación femenina de 1950 y el cambio social (1950-2000)
von: Manuel Pérez Rúa
Veröffentlicht: (2013)
von: Manuel Pérez Rúa
Veröffentlicht: (2013)
Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
von: Liu, Haozhe, et al.
Veröffentlicht: (2025)
von: Liu, Haozhe, et al.
Veröffentlicht: (2025)
SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention
von: Csordás, Róbert, et al.
Veröffentlicht: (2023)
von: Csordás, Róbert, et al.
Veröffentlicht: (2023)
OptMark: Robust Multi-bit Diffusion Watermarking via Inference Time Optimization
von: Xing, Jiazheng, et al.
Veröffentlicht: (2025)
von: Xing, Jiazheng, et al.
Veröffentlicht: (2025)
Learning Video Context as Interleaved Multimodal Sequences
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
Factorized Learning for Temporally Grounded Video-Language Models
von: Zeng, Wenzheng, et al.
Veröffentlicht: (2025)
von: Zeng, Wenzheng, et al.
Veröffentlicht: (2025)
Mitty: Diffusion-based Human-to-Robot Video Generation
von: Song, Yiren, et al.
Veröffentlicht: (2025)
von: Song, Yiren, et al.
Veröffentlicht: (2025)
OmniPSD: Layered PSD Generation with Diffusion Transformer
von: Liu, Cheng, et al.
Veröffentlicht: (2025)
von: Liu, Cheng, et al.
Veröffentlicht: (2025)
Automated Movie Generation via Multi-Agent CoT Planning
von: Wu, Weijia, et al.
Veröffentlicht: (2025)
von: Wu, Weijia, et al.
Veröffentlicht: (2025)
PrimeComposer: Faster Progressively Combined Diffusion for Image Composition with Attention Steering
von: Wang, Yibin, et al.
Veröffentlicht: (2024)
von: Wang, Yibin, et al.
Veröffentlicht: (2024)
Measuring In-Context Computation Complexity via Hidden State Prediction
von: Herrmann, Vincent, et al.
Veröffentlicht: (2025)
von: Herrmann, Vincent, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Learning Useful Representations of Recurrent Neural Network Weight Matrices
von: Herrmann, Vincent, et al.
Veröffentlicht: (2024) -
Lazy Layers to Make Fine-Tuned Diffusion Models More Traceable
von: Liu, Haozhe, et al.
Veröffentlicht: (2024) -
Show-o2: Improved Native Unified Multimodal Models
von: Xie, Jinheng, et al.
Veröffentlicht: (2025) -
WMAdapter: Adding WaterMark Control to Latent Diffusion Models
von: Ci, Hai, et al.
Veröffentlicht: (2024) -
MarDini: Masked Autoregressive Diffusion for Video Generation at Scale
von: Liu, Haozhe, et al.
Veröffentlicht: (2024)