Towards Holistic Modeling for Video Frame Interpolation with Auto-regressive Diffusion Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Peng, Xinyu, Li, Han, Huang, Yuyang, Zheng, Ziyang, Wang, Yaoming, Chen, Xin, Dai, Wenrui, Li, Chenglin, Zou, Junni, Xiong, Hongkai |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HiPART: Hierarchical Pose AutoRegressive Transformer for Occluded 3D Human Pose Estimation
by: Zheng, Hongwei, et al.
Published: (2025)
by: Zheng, Hongwei, et al.
Published: (2025)
Improving Diffusion Models for Inverse Problems Using Optimal Posterior Covariance
by: Peng, Xinyu, et al.
Published: (2024)
by: Peng, Xinyu, et al.
Published: (2024)
Noise Conditional Variational Score Distillation
by: Peng, Xinyu, et al.
Published: (2025)
by: Peng, Xinyu, et al.
Published: (2025)
Frequency-Aware Transformer for Learned Image Compression
by: Li, Han, et al.
Published: (2023)
by: Li, Han, et al.
Published: (2023)
Point Cloud Resampling with Learnable Heat Diffusion
by: Xu, Wenqiang, et al.
Published: (2024)
by: Xu, Wenqiang, et al.
Published: (2024)
Information-Theoretic Optimization for Task-Adapted Compressed Sensing Magnetic Resonance Imaging
by: Peng, Xinyu, et al.
Published: (2026)
by: Peng, Xinyu, et al.
Published: (2026)
3DGabSplat: 3D Gabor Splatting for Frequency-adaptive Radiance Field Rendering
by: Zhou, Junyu, et al.
Published: (2025)
by: Zhou, Junyu, et al.
Published: (2025)
MimicNorm: Weight Mean and Last BN Layer Mimic the Dynamic of Batch Normalization
by: Fei, Wen, et al.
Published: (2020)
by: Fei, Wen, et al.
Published: (2020)
Point Cloud Denoising With Fine-Granularity Dynamic Graph Convolutional Networks
by: Xu, Wenqiang, et al.
Published: (2024)
by: Xu, Wenqiang, et al.
Published: (2024)
On Disentangled Training for Nonlinear Transform in Learned Image Compression
by: Li, Han, et al.
Published: (2025)
by: Li, Han, et al.
Published: (2025)
Error-Propagation-Free Learned Video Compression With Dual-Domain Progressive Temporal Alignment
by: Li, Han, et al.
Published: (2025)
by: Li, Han, et al.
Published: (2025)
Image Compression for Machine and Human Vision with Spatial-Frequency Adaptation
by: Li, Han, et al.
Published: (2024)
by: Li, Han, et al.
Published: (2024)
OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and Generation
by: Li, Han, et al.
Published: (2025)
by: Li, Han, et al.
Published: (2025)
Diffusion-Driven Progressive Target Manipulation for Source-Free Domain Adaptation
by: Huang, Yuyang, et al.
Published: (2025)
by: Huang, Yuyang, et al.
Published: (2025)
METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models
by: Liu, Yuchen, et al.
Published: (2025)
by: Liu, Yuchen, et al.
Published: (2025)
UMG-CLIP: A Unified Multi-Granularity Vision Generalist for Open-World Understanding
by: Shi, Bowen, et al.
Published: (2024)
by: Shi, Bowen, et al.
Published: (2024)
Motion-Aware Video Frame Interpolation
by: Han, Pengfei, et al.
Published: (2024)
by: Han, Pengfei, et al.
Published: (2024)
Video Frame Interpolation for Polarization via Swin-Transformer
by: Huang, Feng, et al.
Published: (2024)
by: Huang, Feng, et al.
Published: (2024)
EventDiff: A Unified and Efficient Diffusion Model Framework for Event-based Video Frame Interpolation
by: Zheng, Hanle, et al.
Published: (2025)
by: Zheng, Hanle, et al.
Published: (2025)
GranViT: A Fine-Grained Vision Model With Autoregressive Perception For MLLMs
by: Zheng, Guanghao, et al.
Published: (2025)
by: Zheng, Guanghao, et al.
Published: (2025)
Perception-Oriented Video Frame Interpolation via Asymmetric Blending
by: Wu, Guangyang, et al.
Published: (2024)
by: Wu, Guangyang, et al.
Published: (2024)
Motion-aware Latent Diffusion Models for Video Frame Interpolation
by: Huang, Zhilin, et al.
Published: (2024)
by: Huang, Zhilin, et al.
Published: (2024)
EV-NVC: Efficient Variable bitrate Neural Video Compression
by: Hu, Yongcun, et al.
Published: (2025)
by: Hu, Yongcun, et al.
Published: (2025)
Can Video Diffusion Models Predict Past Frames? Bidirectional Cycle Consistency for Reversible Interpolation
by: Liu, Lingyu, et al.
Published: (2026)
by: Liu, Lingyu, et al.
Published: (2026)
LDMVFI: Video Frame Interpolation with Latent Diffusion Models
by: Danier, Duolikun, et al.
Published: (2023)
by: Danier, Duolikun, et al.
Published: (2023)
Capturing Conditional Dependence via Auto-regressive Diffusion Models
by: Huang, Xunpeng, et al.
Published: (2025)
by: Huang, Xunpeng, et al.
Published: (2025)
Frame Interpolation with Consecutive Brownian Bridge Diffusion
by: Lyu, Zonglin, et al.
Published: (2024)
by: Lyu, Zonglin, et al.
Published: (2024)
LADDER: An Efficient Framework for Video Frame Interpolation
by: Shen, Tong, et al.
Published: (2024)
by: Shen, Tong, et al.
Published: (2024)
Time-adaptive Video Frame Interpolation based on Residual Diffusion
by: Chavez, Victor Fonte, et al.
Published: (2025)
by: Chavez, Victor Fonte, et al.
Published: (2025)
Benchmarking Video Frame Interpolation
by: Kiefhaber, Simon, et al.
Published: (2024)
by: Kiefhaber, Simon, et al.
Published: (2024)
Generalizable Implicit Motion Modeling for Video Frame Interpolation
by: Guo, Zujin, et al.
Published: (2024)
by: Guo, Zujin, et al.
Published: (2024)
Unified Arbitrary-Time Video Frame Interpolation and Prediction
by: Jin, Xin, et al.
Published: (2025)
by: Jin, Xin, et al.
Published: (2025)
Adapting Image-to-Video Diffusion Models for Large-Motion Frame Interpolation
by: Jin, Luoxu, et al.
Published: (2024)
by: Jin, Luoxu, et al.
Published: (2024)
Velocity Disambiguation for Video Frame Interpolation
by: Zhong, Zhihang, et al.
Published: (2023)
by: Zhong, Zhihang, et al.
Published: (2023)
Beyond Boundary Frames: Context-Centric Video Interpolation with Audio-Visual Semantics
by: Deng, Yuchen, et al.
Published: (2025)
by: Deng, Yuchen, et al.
Published: (2025)
EGVD: Event-Guided Video Diffusion Model for Physically Realistic Large-Motion Frame Interpolation
by: Zhang, Ziran, et al.
Published: (2025)
by: Zhang, Ziran, et al.
Published: (2025)
Network-based Neighborhood regression
by: Zhen, Yaoming, et al.
Published: (2024)
by: Zhen, Yaoming, et al.
Published: (2024)
VTinker: Guided Flow Upsampling and Texture Mapping for High-Resolution Video Frame Interpolation
by: Wu, Chenyang, et al.
Published: (2025)
by: Wu, Chenyang, et al.
Published: (2025)
EDEN: Enhanced Diffusion for High-quality Large-motion Video Frame Interpolation
by: Zhang, Zihao, et al.
Published: (2025)
by: Zhang, Zihao, et al.
Published: (2025)
Disentangled Motion Modeling for Video Frame Interpolation
by: Lew, Jaihyun, et al.
Published: (2024)
by: Lew, Jaihyun, et al.
Published: (2024)
Similar Items
-
HiPART: Hierarchical Pose AutoRegressive Transformer for Occluded 3D Human Pose Estimation
by: Zheng, Hongwei, et al.
Published: (2025) -
Improving Diffusion Models for Inverse Problems Using Optimal Posterior Covariance
by: Peng, Xinyu, et al.
Published: (2024) -
Noise Conditional Variational Score Distillation
by: Peng, Xinyu, et al.
Published: (2025) -
Frequency-Aware Transformer for Learned Image Compression
by: Li, Han, et al.
Published: (2023) -
Point Cloud Resampling with Learnable Heat Diffusion
by: Xu, Wenqiang, et al.
Published: (2024)