Saved in:
| Main Authors: | Wen, Yuxin, Wu, Jim, Jain, Ajay, Goldstein, Tom, Panda, Ashwinee |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2504.10317 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FineGRAIN: Evaluating Failure Modes of Text-to-Image Models with Vision Language Model Judges
by: Hayes, Kevin David, et al.
Published: (2025)
by: Hayes, Kevin David, et al.
Published: (2025)
Attention Sinks in Diffusion Transformers: A Causal Analysis
by: Wu, Fangzheng, et al.
Published: (2026)
by: Wu, Fangzheng, et al.
Published: (2026)
Bidirectional Sparse Attention for Faster Video Diffusion Training
by: Zhan, Chenlu, et al.
Published: (2025)
by: Zhan, Chenlu, et al.
Published: (2025)
VMonarch: Efficient Video Diffusion Transformers with Structured Attention
by: Liang, Cheng, et al.
Published: (2026)
by: Liang, Cheng, et al.
Published: (2026)
EditScout: Locating Forged Regions from Diffusion-based Edited Images with Multimodal LLM
by: Nguyen, Quang, et al.
Published: (2024)
by: Nguyen, Quang, et al.
Published: (2024)
ReHyAt: Recurrent Hybrid Attention for Video Diffusion Transformers
by: Ghafoorian, Mohsen, et al.
Published: (2026)
by: Ghafoorian, Mohsen, et al.
Published: (2026)
Quantifying Cross-Modality Memorization in Vision-Language Models
by: Wen, Yuxin, et al.
Published: (2025)
by: Wen, Yuxin, et al.
Published: (2025)
QuantSparse: Comprehensively Compressing Video Diffusion Transformer with Model Quantization and Attention Sparsification
by: Feng, Weilun, et al.
Published: (2025)
by: Feng, Weilun, et al.
Published: (2025)
ARGUS: Hallucination and Omission Evaluation in Video-LLMs
by: Rawal, Ruchit, et al.
Published: (2025)
by: Rawal, Ruchit, et al.
Published: (2025)
Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformer
by: Ghafoorian, Mohsen, et al.
Published: (2025)
by: Ghafoorian, Mohsen, et al.
Published: (2025)
Efficient-vDiT: Efficient Video Diffusion Transformers With Attention Tile
by: Ding, Hangliang, et al.
Published: (2025)
by: Ding, Hangliang, et al.
Published: (2025)
FrameDiT: Diffusion Transformer with Matrix Attention for Efficient Video Generation
by: Le, Minh Khoa, et al.
Published: (2026)
by: Le, Minh Khoa, et al.
Published: (2026)
One-Shot Learning Meets Depth Diffusion in Multi-Object Videos
by: Jain, Anisha
Published: (2024)
by: Jain, Anisha
Published: (2024)
SALAD: Achieve High-Sparsity Attention via Efficient Linear Attention Tuning for Video Diffusion Transformer
by: Fang, Tongcheng, et al.
Published: (2026)
by: Fang, Tongcheng, et al.
Published: (2026)
DiTPainter: Efficient Video Inpainting with Diffusion Transformers
by: Wu, Xian, et al.
Published: (2025)
by: Wu, Xian, et al.
Published: (2025)
LiteAttention: A Temporal Sparse Attention for Diffusion Transformers
by: Shmilovich, Dor, et al.
Published: (2025)
by: Shmilovich, Dor, et al.
Published: (2025)
VMoBA: Mixture-of-Block Attention for Video Diffusion Models
by: Wu, Jianzong, et al.
Published: (2025)
by: Wu, Jianzong, et al.
Published: (2025)
AttentionBender: Manipulating Cross-Attention in Video Diffusion Transformers as a Creative Probe
by: Cole, Adam, et al.
Published: (2026)
by: Cole, Adam, et al.
Published: (2026)
Video Interpolation with Diffusion Models
by: Jain, Siddhant, et al.
Published: (2024)
by: Jain, Siddhant, et al.
Published: (2024)
Understanding Attention Mechanism in Video Diffusion Models
by: Liu, Bingyan, et al.
Published: (2025)
by: Liu, Bingyan, et al.
Published: (2025)
Blended Latent Diffusion under Attention Control for Real-World Video Editing
by: Liu, Deyin, et al.
Published: (2024)
by: Liu, Deyin, et al.
Published: (2024)
Perceiving Longer Sequences With Bi-Directional Cross-Attention Transformers
by: Hiller, Markus, et al.
Published: (2024)
by: Hiller, Markus, et al.
Published: (2024)
CustomVideoX: 3D Reference Attention Driven Dynamic Adaptation for Zero-Shot Customized Video Diffusion Transformers
by: She, D., et al.
Published: (2025)
by: She, D., et al.
Published: (2025)
DiTVR: Zero-Shot Diffusion Transformer for Video Restoration
by: Gao, Sicheng, et al.
Published: (2025)
by: Gao, Sicheng, et al.
Published: (2025)
BinaryAttention: One-Bit QK-Attention for Vision and Diffusion Transformers
by: Xiao, Chaodong, et al.
Published: (2026)
by: Xiao, Chaodong, et al.
Published: (2026)
Mixture of Distributions Matters: Dynamic Sparse Attention for Efficient Video Diffusion Transformers
by: Liu, Yuxi, et al.
Published: (2026)
by: Liu, Yuxi, et al.
Published: (2026)
Hierarchical Point Attention for Indoor 3D Object Detection
by: Shu, Manli, et al.
Published: (2023)
by: Shu, Manli, et al.
Published: (2023)
Video Diffusion Transformers are In-Context Learners
by: Fei, Zhengcong, et al.
Published: (2024)
by: Fei, Zhengcong, et al.
Published: (2024)
VSA: Faster Video Diffusion with Trainable Sparse Attention
by: Zhang, Peiyuan, et al.
Published: (2025)
by: Zhang, Peiyuan, et al.
Published: (2025)
Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation
by: Wang, Wenjing, et al.
Published: (2023)
by: Wang, Wenjing, et al.
Published: (2023)
Detecting, Explaining, and Mitigating Memorization in Diffusion Models
by: Wen, Yuxin, et al.
Published: (2024)
by: Wen, Yuxin, et al.
Published: (2024)
SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers
by: Qiu, Di, et al.
Published: (2025)
by: Qiu, Di, et al.
Published: (2025)
DVD-Quant: Data-free Video Diffusion Transformers Quantization
by: Li, Zhiteng, et al.
Published: (2025)
by: Li, Zhiteng, et al.
Published: (2025)
Decouple and Track: Benchmarking and Improving Video Diffusion Transformers for Motion Transfer
by: Shi, Qingyu, et al.
Published: (2025)
by: Shi, Qingyu, et al.
Published: (2025)
Image Generation with a Sphere Encoder
by: Yue, Kaiyu, et al.
Published: (2026)
by: Yue, Kaiyu, et al.
Published: (2026)
In-Context Audio Control of Video Diffusion Transformers
by: Liu, Wenze, et al.
Published: (2025)
by: Liu, Wenze, et al.
Published: (2025)
Latte: Latent Diffusion Transformer for Video Generation
by: Ma, Xin, et al.
Published: (2024)
by: Ma, Xin, et al.
Published: (2024)
Steering Video Diffusion Transformers with Massive Activations
by: Cheng, Xianhang, et al.
Published: (2026)
by: Cheng, Xianhang, et al.
Published: (2026)
VORTA: Efficient Video Diffusion via Routing Sparse Attention
by: Sun, Wenhao, et al.
Published: (2025)
by: Sun, Wenhao, et al.
Published: (2025)
Veda: Scalable Video Diffusion via Distilled Sparse Attention
by: Han, Shihao, et al.
Published: (2026)
by: Han, Shihao, et al.
Published: (2026)
Similar Items
-
FineGRAIN: Evaluating Failure Modes of Text-to-Image Models with Vision Language Model Judges
by: Hayes, Kevin David, et al.
Published: (2025) -
Attention Sinks in Diffusion Transformers: A Causal Analysis
by: Wu, Fangzheng, et al.
Published: (2026) -
Bidirectional Sparse Attention for Faster Video Diffusion Training
by: Zhan, Chenlu, et al.
Published: (2025) -
VMonarch: Efficient Video Diffusion Transformers with Structured Attention
by: Liang, Cheng, et al.
Published: (2026) -
EditScout: Locating Forged Regions from Diffusion-based Edited Images with Multimodal LLM
by: Nguyen, Quang, et al.
Published: (2024)