Flow and Depth Assisted Video Prediction with Latent Transformer
Fuente:
arXiv
Saved in:
| Main Authors: | Suleyman, Eliyas, Henderson, Paul, Firkat, Eksan, Pugeault, Nicolas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Benefits of Instance Decomposition in Video Prediction Models
by: Suleyman, Eliyas, et al.
Published: (2025)
by: Suleyman, Eliyas, et al.
Published: (2025)
Beyond Reconstruction: A Physics Based Neural Deferred Shader for Photo-realistic Rendering
by: He, Zhuo, et al.
Published: (2025)
by: He, Zhuo, et al.
Published: (2025)
Generative Fields: Uncovering Hierarchical Feature Control for StyleGAN via Inverted Receptive Fields
by: He, Zhuo, et al.
Published: (2025)
by: He, Zhuo, et al.
Published: (2025)
A Convolutional Neural Deferred Shader for Physics Based Rendering
by: He, Zhuo, et al.
Published: (2025)
by: He, Zhuo, et al.
Published: (2025)
Splat-Portrait: Generalizing Talking Heads with Gaussian Splatting
by: Shi, Tong, et al.
Published: (2026)
by: Shi, Tong, et al.
Published: (2026)
Detail-Enhanced Intra- and Inter-modal Interaction for Audio-Visual Emotion Recognition
by: Shi, Tong, et al.
Published: (2024)
by: Shi, Tong, et al.
Published: (2024)
The Bad Batches: Enhancing Self-Supervised Learning in Image Classification Through Representative Batch Curation
by: Goksu, Ozgu, et al.
Published: (2024)
by: Goksu, Ozgu, et al.
Published: (2024)
FedQuad: Federated Stochastic Quadruplet Learning to Mitigate Data Heterogeneity
by: Goksu, Ozgu, et al.
Published: (2025)
by: Goksu, Ozgu, et al.
Published: (2025)
Enhancing Federated Quadruplet Learning: Stochastic Client Selection and Embedding Stability Analysis
by: Goksu, Ozgu, et al.
Published: (2026)
by: Goksu, Ozgu, et al.
Published: (2026)
Hybrid-Regularized Magnitude Pruning for Robust Federated Learning under Covariate Shift
by: Goksu, Ozgu, et al.
Published: (2024)
by: Goksu, Ozgu, et al.
Published: (2024)
DepthFlow: Exploiting Depth-Flow Structural Correlations for Unsupervised Video Object Segmentation
by: Cho, Suhwan, et al.
Published: (2025)
by: Cho, Suhwan, et al.
Published: (2025)
Video Generation with Predictive Latents
by: Zhao, Yian, et al.
Published: (2026)
by: Zhao, Yian, et al.
Published: (2026)
FutureDepth: Learning to Predict the Future Improves Video Depth Estimation
by: Yasarla, Rajeev, et al.
Published: (2024)
by: Yasarla, Rajeev, et al.
Published: (2024)
Latte: Latent Diffusion Transformer for Video Generation
by: Ma, Xin, et al.
Published: (2024)
by: Ma, Xin, et al.
Published: (2024)
Modeling 3D Pedestrian-Vehicle Interactions for Vehicle-Conditioned Pose Forecasting
by: Zhu, Guangxun, et al.
Published: (2026)
by: Zhu, Guangxun, et al.
Published: (2026)
Simplify Implant Depth Prediction as Video Grounding: A Texture Perceive Implant Depth Prediction Network
by: Yang, Xinquan, et al.
Published: (2024)
by: Yang, Xinquan, et al.
Published: (2024)
A Transformer-Based Model for the Prediction of Human Gaze Behavior on Videos
by: Ozdel, Suleyman, et al.
Published: (2024)
by: Ozdel, Suleyman, et al.
Published: (2024)
Online Video Depth Anything: Temporally-Consistent Depth Prediction with Low Memory Consumption
by: Feiden, Johann-Friedrich, et al.
Published: (2025)
by: Feiden, Johann-Friedrich, et al.
Published: (2025)
Virtually Unrolling the Herculaneum Papyri by Diffeomorphic Spiral Fitting
by: Henderson, Paul
Published: (2025)
by: Henderson, Paul
Published: (2025)
IDOL: Unified Dual-Modal Latent Diffusion for Human-Centric Joint Video-Depth Generation
by: Zhai, Yuanhao, et al.
Published: (2024)
by: Zhai, Yuanhao, et al.
Published: (2024)
Efficient Continuous Video Flow Model for Video Prediction
by: Shrivastava, Gaurav, et al.
Published: (2024)
by: Shrivastava, Gaurav, et al.
Published: (2024)
VEDIT: Latent Prediction Architecture For Procedural Video Representation Learning
by: Lin, Han, et al.
Published: (2024)
by: Lin, Han, et al.
Published: (2024)
DaBiT: Depth and Blur informed Transformer for Video Focal Deblurring
by: Morris, Crispian, et al.
Published: (2024)
by: Morris, Crispian, et al.
Published: (2024)
Seer: Language Instructed Video Prediction with Latent Diffusion Models
by: Gu, Xianfan, et al.
Published: (2023)
by: Gu, Xianfan, et al.
Published: (2023)
JLT: Clean-Latent Prediction in Latent Diffusion Transformers
by: Fu, Funing, et al.
Published: (2026)
by: Fu, Funing, et al.
Published: (2026)
FlashDepth: Real-time Streaming Video Depth Estimation at 2K Resolution
by: Chou, Gene, et al.
Published: (2025)
by: Chou, Gene, et al.
Published: (2025)
Video Depth Propagation
by: Piccinelli, Luigi, et al.
Published: (2025)
by: Piccinelli, Luigi, et al.
Published: (2025)
Video Depth without Video Models
by: Ke, Bingxin, et al.
Published: (2024)
by: Ke, Bingxin, et al.
Published: (2024)
RD-ViT: Recurrent-Depth Vision Transformer for Semantic Segmentation with Reduced Data Dependence Extending the Recurrent-Depth Transformer Architecture to Dense Prediction
by: He, Renjie
Published: (2026)
by: He, Renjie
Published: (2026)
AssistPDA: An Online Video Surveillance Assistant for Video Anomaly Prediction, Detection, and Analysis
by: Yang, Zhiwei, et al.
Published: (2025)
by: Yang, Zhiwei, et al.
Published: (2025)
DepthFM: Fast Monocular Depth Estimation with Flow Matching
by: Gui, Ming, et al.
Published: (2024)
by: Gui, Ming, et al.
Published: (2024)
Video Prediction Transformers without Recurrence or Convolution
by: Tang, Yujin, et al.
Published: (2024)
by: Tang, Yujin, et al.
Published: (2024)
Exploiting Optical Flow Guidance for Transformer-Based Video Inpainting
by: Zhang, Kaidong, et al.
Published: (2023)
by: Zhang, Kaidong, et al.
Published: (2023)
Inference-time Physics Alignment of Video Generative Models with Latent World Models
by: Yuan, Jianhao, et al.
Published: (2026)
by: Yuan, Jianhao, et al.
Published: (2026)
Sampling 3D Gaussian Scenes in Seconds with Latent Diffusion Models
by: Henderson, Paul, et al.
Published: (2024)
by: Henderson, Paul, et al.
Published: (2024)
Depth-Wise Representation Development Under Blockwise Self-Supervised Learning for Video Vision Transformers
by: Römer, Jonas, et al.
Published: (2026)
by: Römer, Jonas, et al.
Published: (2026)
LTX-Video: Realtime Video Latent Diffusion
by: HaCohen, Yoav, et al.
Published: (2024)
by: HaCohen, Yoav, et al.
Published: (2024)
Exploiting Style Latent Flows for Generalizing Deepfake Video Detection
by: Choi, Jongwook, et al.
Published: (2024)
by: Choi, Jongwook, et al.
Published: (2024)
DeVOS: Flow-Guided Deformable Transformer for Video Object Segmentation
by: Fedynyak, Volodymyr, et al.
Published: (2024)
by: Fedynyak, Volodymyr, et al.
Published: (2024)
Latent Video Dataset Distillation
by: Li, Ning, et al.
Published: (2025)
by: Li, Ning, et al.
Published: (2025)
Similar Items
-
On the Benefits of Instance Decomposition in Video Prediction Models
by: Suleyman, Eliyas, et al.
Published: (2025) -
Beyond Reconstruction: A Physics Based Neural Deferred Shader for Photo-realistic Rendering
by: He, Zhuo, et al.
Published: (2025) -
Generative Fields: Uncovering Hierarchical Feature Control for StyleGAN via Inverted Receptive Fields
by: He, Zhuo, et al.
Published: (2025) -
A Convolutional Neural Deferred Shader for Physics Based Rendering
by: He, Zhuo, et al.
Published: (2025) -
Splat-Portrait: Generalizing Talking Heads with Gaussian Splatting
by: Shi, Tong, et al.
Published: (2026)