ReVision: Refining Video Diffusion with Explicit 3D Motion Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Qihao, He, Ju, Yu, Qihang, Chen, Liang-Chieh, Yuille, Alan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers
by: Ren, Sucheng, et al.
Published: (2025)
by: Ren, Sucheng, et al.
Published: (2025)
ViTamin: Designing Scalable Vision Models in the Vision-Language Era
by: Chen, Jieneng, et al.
Published: (2024)
by: Chen, Jieneng, et al.
Published: (2024)
FlowTok: Flowing Seamlessly Across Text and Image Tokens
by: He, Ju, et al.
Published: (2025)
by: He, Ju, et al.
Published: (2025)
Alleviating Distortion in Image Generation via Multi-Resolution Diffusion Models and Time-Dependent Layer Normalization
by: Liu, Qihao, et al.
Published: (2024)
by: Liu, Qihao, et al.
Published: (2024)
A Simple Video Segmenter by Tracking Objects Along Axial Trajectories
by: He, Ju, et al.
Published: (2023)
by: He, Ju, et al.
Published: (2023)
Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation
by: Ren, Sucheng, et al.
Published: (2025)
by: Ren, Sucheng, et al.
Published: (2025)
Frequency-Aware Flow Matching for High-Quality Image Generation
by: Ren, Sucheng, et al.
Published: (2026)
by: Ren, Sucheng, et al.
Published: (2026)
FlowAR: Scale-wise Autoregressive Image Generation Meets Flow Matching
by: Ren, Sucheng, et al.
Published: (2024)
by: Ren, Sucheng, et al.
Published: (2024)
Dictionary-based Framework for Interpretable and Consistent Object Parsing
by: Zhang, Tiezheng, et al.
Published: (2025)
by: Zhang, Tiezheng, et al.
Published: (2025)
Autoregressive Image Generation with Masked Bit Modeling
by: Yu, Qihang, et al.
Published: (2026)
by: Yu, Qihang, et al.
Published: (2026)
SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
by: Ma, Wufei, et al.
Published: (2025)
by: Ma, Wufei, et al.
Published: (2025)
TriDiff-4D: Fast 4D Generation through Diffusion-based Triplane Re-posing
by: Sheung, Eddie Pokming, et al.
Published: (2025)
by: Sheung, Eddie Pokming, et al.
Published: (2025)
SPFormer: Enhancing Vision Transformer with Superpixel Representation
by: Mei, Jieru, et al.
Published: (2024)
by: Mei, Jieru, et al.
Published: (2024)
DIRECT-3D: Learning Direct Text-to-3D Generation on Massive Noisy 3D Data
by: Liu, Qihao, et al.
Published: (2024)
by: Liu, Qihao, et al.
Published: (2024)
World-consistent Video Diffusion with Explicit 3D Modeling
by: Zhang, Qihang, et al.
Published: (2024)
by: Zhang, Qihang, et al.
Published: (2024)
4D-Animal: Freely Reconstructing Animatable 3D Animals from Videos
by: Zhong, Shanshan, et al.
Published: (2025)
by: Zhong, Shanshan, et al.
Published: (2025)
Explicit Critic Guidance for Aligning Diffusion Models
by: Liang, Zhengyang, et al.
Published: (2026)
by: Liang, Zhengyang, et al.
Published: (2026)
Enhancing Temporal Consistency in Video Editing by Reconstructing Videos with 3D Gaussian Splatting
by: Shin, Inkyu, et al.
Published: (2024)
by: Shin, Inkyu, et al.
Published: (2024)
Randomized Autoregressive Visual Generation
by: Yu, Qihang, et al.
Published: (2024)
by: Yu, Qihang, et al.
Published: (2024)
Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models
by: Zhang, Tiezheng, et al.
Published: (2025)
by: Zhang, Tiezheng, et al.
Published: (2025)
ImageNet3D: Towards General-Purpose Object-Level 3D Understanding
by: Ma, Wufei, et al.
Published: (2024)
by: Ma, Wufei, et al.
Published: (2024)
Generating Images with 3D Annotations Using Diffusion Models
by: Ma, Wufei, et al.
Published: (2023)
by: Ma, Wufei, et al.
Published: (2023)
Can These Views Be One Scene? Evaluating Multiview 3D Consistency when 3D Foundation Models Hallucinate
by: Paul, Soumava, et al.
Published: (2026)
by: Paul, Soumava, et al.
Published: (2026)
Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution
by: Liu, Qihao, et al.
Published: (2024)
by: Liu, Qihao, et al.
Published: (2024)
Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens
by: Kim, Dongwon, et al.
Published: (2025)
by: Kim, Dongwon, et al.
Published: (2025)
Differences That Matter: Auditing Models for Capability Gap Discovery and Rectification
by: Liu, Qihao, et al.
Published: (2025)
by: Liu, Qihao, et al.
Published: (2025)
Computer Vision and Its Relationship to Cognitive Science: A perspective from Bayes Decision Theory
by: Yuille, Alan, et al.
Published: (2026)
by: Yuille, Alan, et al.
Published: (2026)
Rethinking Video-Text Understanding: Retrieval from Counterfactually Augmented Data
by: Ma, Wufei, et al.
Published: (2024)
by: Ma, Wufei, et al.
Published: (2024)
ReVision: A Dataset and Baseline VLM for Privacy-Preserving Task-Oriented Visual Instruction Rewriting
by: Mishra, Abhijit, et al.
Published: (2025)
by: Mishra, Abhijit, et al.
Published: (2025)
RePack then Refine: Efficient Diffusion Transformer with Vision Foundation Model
by: Dong, Guanfang, et al.
Published: (2025)
by: Dong, Guanfang, et al.
Published: (2025)
EvoWorld: Evolving Panoramic World Generation with Explicit 3D Memory
by: Wang, Jiahao, et al.
Published: (2025)
by: Wang, Jiahao, et al.
Published: (2025)
DINeMo: Learning Neural Mesh Models with no 3D Annotations
by: Guo, Weijie, et al.
Published: (2025)
by: Guo, Weijie, et al.
Published: (2025)
Exploring Iterative Refinement with Diffusion Models for Video Grounding
by: Liang, Xiao, et al.
Published: (2023)
by: Liang, Xiao, et al.
Published: (2023)
A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens
by: Kerssies, Tommie, et al.
Published: (2026)
by: Kerssies, Tommie, et al.
Published: (2026)
SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference
by: Wang, Feng, et al.
Published: (2023)
by: Wang, Feng, et al.
Published: (2023)
DO3D: Self-supervised Learning of Decomposed Object-aware 3D Motion and Depth from Monocular Videos
by: Wu, Xiuzhe, et al.
Published: (2024)
by: Wu, Xiuzhe, et al.
Published: (2024)
Efficient Large Multi-modal Models via Visual Context Compression
by: Chen, Jieneng, et al.
Published: (2024)
by: Chen, Jieneng, et al.
Published: (2024)
Large Language Models are Universal Reasoners for Visual Generation
by: Ren, Sucheng, et al.
Published: (2026)
by: Ren, Sucheng, et al.
Published: (2026)
Motion Diffusion-Guided 3D Global HMR from a Dynamic Camera
by: Heo, Jaewoo, et al.
Published: (2024)
by: Heo, Jaewoo, et al.
Published: (2024)
Gaussian Scenes: Pose-Free Sparse-View Scene Reconstruction using Depth-Enhanced Diffusion Priors
by: Paul, Soumava, et al.
Published: (2024)
by: Paul, Soumava, et al.
Published: (2024)
Similar Items
-
Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers
by: Ren, Sucheng, et al.
Published: (2025) -
ViTamin: Designing Scalable Vision Models in the Vision-Language Era
by: Chen, Jieneng, et al.
Published: (2024) -
FlowTok: Flowing Seamlessly Across Text and Image Tokens
by: He, Ju, et al.
Published: (2025) -
Alleviating Distortion in Image Generation via Multi-Resolution Diffusion Models and Time-Dependent Layer Normalization
by: Liu, Qihao, et al.
Published: (2024) -
A Simple Video Segmenter by Tracking Objects Along Axial Trajectories
by: He, Ju, et al.
Published: (2023)