Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Samuel, Dvir, Tzachor, Issar, Levy, Matan, Green, Micahel, Chechik, Gal, Ben-Ari, Rami |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval
by: Tzachor, Issar, et al.
Published: (2026)
by: Tzachor, Issar, et al.
Published: (2026)
Find your Needle: Small Object Image Retrieval via Multi-Object Attention Optimization
by: Green, Michael, et al.
Published: (2025)
by: Green, Michael, et al.
Published: (2025)
OmnimatteZero: Fast Training-free Omnimatte with Pre-trained Video Diffusion Models
by: Samuel, Dvir, et al.
Published: (2025)
by: Samuel, Dvir, et al.
Published: (2025)
Where's Waldo: Diffusion Features for Personalized Segmentation and Retrieval
by: Samuel, Dvir, et al.
Published: (2024)
by: Samuel, Dvir, et al.
Published: (2024)
Retrieval-Augmented Gaussian Avatars: Improving Expression Generalization
by: Levy, Matan, et al.
Published: (2026)
by: Levy, Matan, et al.
Published: (2026)
Fast 4D Mesh Generation by Spatio-Temporal Attention Chains
by: Samuel, Dvir, et al.
Published: (2026)
by: Samuel, Dvir, et al.
Published: (2026)
EffoVPR: Effective Foundation Model Utilization for Visual Place Recognition
by: Tzachor, Issar, et al.
Published: (2024)
by: Tzachor, Issar, et al.
Published: (2024)
Lightning-Fast Image Inversion and Editing for Text-to-Image Diffusion Models
by: Samuel, Dvir, et al.
Published: (2023)
by: Samuel, Dvir, et al.
Published: (2023)
Task-Specific Adaptation with Restricted Model Access
by: Levy, Matan, et al.
Published: (2025)
by: Levy, Matan, et al.
Published: (2025)
Per-Query Visual Concept Learning
by: Malca, Ori, et al.
Published: (2025)
by: Malca, Ori, et al.
Published: (2025)
Set Features for Anomaly Detection
by: Cohen, Niv, et al.
Published: (2023)
by: Cohen, Niv, et al.
Published: (2023)
Bringing Objects to Life: training-free 4D generation from 3D objects through view consistent noise
by: Rahamim, Ohad, et al.
Published: (2024)
by: Rahamim, Ohad, et al.
Published: (2024)
Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion Models
by: Tewel, Yoad, et al.
Published: (2024)
by: Tewel, Yoad, et al.
Published: (2024)
Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models
by: Ji, Yicheng, et al.
Published: (2026)
by: Ji, Yicheng, et al.
Published: (2026)
Light Forcing: Accelerating Autoregressive Video Diffusion via Sparse Attention
by: Lv, Chengtao, et al.
Published: (2026)
by: Lv, Chengtao, et al.
Published: (2026)
Story2Board: A Training-Free Approach for Expressive Storyboard Generation
by: Dinkevich, David, et al.
Published: (2025)
by: Dinkevich, David, et al.
Published: (2025)
MSC: Multi-Scale Spatio-Temporal Causal Attention for Autoregressive Video Diffusion
by: Xu, Xunnong, et al.
Published: (2024)
by: Xu, Xunnong, et al.
Published: (2024)
Motion by Queries: Identity-Motion Trade-offs in Text-to-Video Generation
by: Atzmon, Yuval, et al.
Published: (2024)
by: Atzmon, Yuval, et al.
Published: (2024)
Sparse Forcing: Native Trainable Sparse Attention for Real-time Autoregressive Diffusion Video Generation
by: Xu, Boxun, et al.
Published: (2026)
by: Xu, Boxun, et al.
Published: (2026)
LiveSVG: Zero-Shot SVG Animation via Video Generation
by: Levy, Matan, et al.
Published: (2026)
by: Levy, Matan, et al.
Published: (2026)
CarGait: Cross-Attention based Re-ranking for Gait recognition
by: Habib, Gavriel, et al.
Published: (2025)
by: Habib, Gavriel, et al.
Published: (2025)
RL-RC-DoT: A Block-level RL agent for Task-Aware Video Compression
by: Gadot, Uri, et al.
Published: (2025)
by: Gadot, Uri, et al.
Published: (2025)
Linguistic Binding in Diffusion Models: Enhancing Attribute Correspondence through Attention Map Alignment
by: Rassin, Royi, et al.
Published: (2023)
by: Rassin, Royi, et al.
Published: (2023)
QuantSparse: Comprehensively Compressing Video Diffusion Transformer with Model Quantization and Attention Sparsification
by: Feng, Weilun, et al.
Published: (2025)
by: Feng, Weilun, et al.
Published: (2025)
Compositional Video Generation via Inference-Time Guidance
by: Shaulov, Ariel, et al.
Published: (2026)
by: Shaulov, Ariel, et al.
Published: (2026)
Past- and Future-Informed KV Cache Policy with Salience Estimation in Autoregressive Video Diffusion
by: Chen, Hanmo, et al.
Published: (2026)
by: Chen, Hanmo, et al.
Published: (2026)
FilmWeaver: Weaving Consistent Multi-Shot Videos with Cache-Guided Autoregressive Diffusion
by: Luo, Xiangyang, et al.
Published: (2025)
by: Luo, Xiangyang, et al.
Published: (2025)
From Slow Bidirectional to Fast Autoregressive Video Diffusion Models
by: Yin, Tianwei, et al.
Published: (2024)
by: Yin, Tianwei, et al.
Published: (2024)
VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion
by: Yesiltepe, Hidir, et al.
Published: (2026)
by: Yesiltepe, Hidir, et al.
Published: (2026)
Head-Aware KV Cache Compression for Efficient Visual Autoregressive Modeling
by: Qin, Ziran, et al.
Published: (2025)
by: Qin, Ziran, et al.
Published: (2025)
Ca2-VDM: Efficient Autoregressive Video Diffusion Model with Causal Generation and Cache Sharing
by: Gao, Kaifeng, et al.
Published: (2024)
by: Gao, Kaifeng, et al.
Published: (2024)
JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion
by: Chen, Anthony, et al.
Published: (2026)
by: Chen, Anthony, et al.
Published: (2026)
LiteAttention: A Temporal Sparse Attention for Diffusion Transformers
by: Shmilovich, Dor, et al.
Published: (2025)
by: Shmilovich, Dor, et al.
Published: (2025)
Data-Driven Loss Functions for Inference-Time Optimization in Text-to-Image
by: Yiflach, Sapir Esther, et al.
Published: (2025)
by: Yiflach, Sapir Esther, et al.
Published: (2025)
X-Cache: Cross-Chunk Block Caching for Few-Step Autoregressive World Models Inference
by: Zeng, Yixiao, et al.
Published: (2026)
by: Zeng, Yixiao, et al.
Published: (2026)
Single Image Iterative Subject-driven Generation and Editing
by: Shpitzer, Yair, et al.
Published: (2025)
by: Shpitzer, Yair, et al.
Published: (2025)
Key-Locked Rank One Editing for Text-to-Image Personalization
by: Tewel, Yoad, et al.
Published: (2023)
by: Tewel, Yoad, et al.
Published: (2023)
Assessing Image Quality Using a Simple Generative Representation
by: Raviv, Simon, et al.
Published: (2024)
by: Raviv, Simon, et al.
Published: (2024)
FastCache: Fast Caching for Diffusion Transformer Through Learnable Linear Approximation
by: Liu, Dong, et al.
Published: (2025)
by: Liu, Dong, et al.
Published: (2025)
MagCache: Fast Video Generation with Magnitude-Aware Cache
by: Ma, Zehong, et al.
Published: (2025)
by: Ma, Zehong, et al.
Published: (2025)
Similar Items
-
VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval
by: Tzachor, Issar, et al.
Published: (2026) -
Find your Needle: Small Object Image Retrieval via Multi-Object Attention Optimization
by: Green, Michael, et al.
Published: (2025) -
OmnimatteZero: Fast Training-free Omnimatte with Pre-trained Video Diffusion Models
by: Samuel, Dvir, et al.
Published: (2025) -
Where's Waldo: Diffusion Features for Personalized Segmentation and Retrieval
by: Samuel, Dvir, et al.
Published: (2024) -
Retrieval-Augmented Gaussian Avatars: Improving Expression Generalization
by: Levy, Matan, et al.
Published: (2026)