Think Before You Diffuse: Infusing Physical Rules into Video Diffusion
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Ke, Xiao, Cihan, Xu, Jiacong, Mei, Yiqun, Patel, Vishal M. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FreeViS: Training-free Video Stylization with Inconsistent References
por: Xu, Jiacong, et al.
Publicado: (2025)
por: Xu, Jiacong, et al.
Publicado: (2025)
Endless World: Real-Time 3D-Aware Long Video Generation
por: Zhang, Ke, et al.
Publicado: (2025)
por: Zhang, Ke, et al.
Publicado: (2025)
Wild-GS: Real-Time Novel View Synthesis from Unconstrained Photo Collections
por: Xu, Jiacong, et al.
Publicado: (2024)
por: Xu, Jiacong, et al.
Publicado: (2024)
Reference-based Controllable Scene Stylization with Gaussian Splatting
por: Mei, Yiqun, et al.
Publicado: (2024)
por: Mei, Yiqun, et al.
Publicado: (2024)
Not All Tokens Need 40 Steps: Heterogeneous Step Allocation in Diffusion Transformers for Efficient Video Generation
por: Chu, Ernie, et al.
Publicado: (2026)
por: Chu, Ernie, et al.
Publicado: (2026)
Frame by Familiar Frame: Understanding Replication in Video Diffusion Models
por: Rahman, Aimon, et al.
Publicado: (2024)
por: Rahman, Aimon, et al.
Publicado: (2024)
Reinforcing Video Reasoning Segmentation to Think Before It Segments
por: Gong, Sitong, et al.
Publicado: (2025)
por: Gong, Sitong, et al.
Publicado: (2025)
Dreamguider: Improved Training free Diffusion-based Conditional Generation
por: Nair, Nithin Gopalakrishnan, et al.
Publicado: (2024)
por: Nair, Nithin Gopalakrishnan, et al.
Publicado: (2024)
Leveraging Thermal Modality to Enhance Reconstruction in Low-Light Conditions
por: Xu, Jiacong, et al.
Publicado: (2024)
por: Xu, Jiacong, et al.
Publicado: (2024)
MedCL: Learning Consistent Anatomy Distribution for Scribble-supervised Medical Image Segmentation
por: Zhang, Ke, et al.
Publicado: (2025)
por: Zhang, Ke, et al.
Publicado: (2025)
ModelMix: A New Model-Mixup Strategy to Minimize Vicinal Risk across Tasks for Few-scribble based Cardiac Segmentation
por: Zhang, Ke, et al.
Publicado: (2024)
por: Zhang, Ke, et al.
Publicado: (2024)
Lux Post Facto: Learning Portrait Performance Relighting with Conditional Video Diffusion and a Hybrid Dataset
por: Mei, Yiqun, et al.
Publicado: (2025)
por: Mei, Yiqun, et al.
Publicado: (2025)
WAT: Online Video Understanding Needs Watching Before Thinking
por: Han, Zifan, et al.
Publicado: (2026)
por: Han, Zifan, et al.
Publicado: (2026)
Learning on the Manifold: Unlocking Standard Diffusion Transformers with Representation Encoders
por: Kumar, Amandeep, et al.
Publicado: (2026)
por: Kumar, Amandeep, et al.
Publicado: (2026)
Your Pre-trained Diffusion Model Secretly Knows Restoration
por: Rajagopalan, Sudarshan, et al.
Publicado: (2026)
por: Rajagopalan, Sudarshan, et al.
Publicado: (2026)
Think Before You Move: Latent Motion Reasoning for Text-to-Motion Generation
por: Qian, Yijie, et al.
Publicado: (2025)
por: Qian, Yijie, et al.
Publicado: (2025)
InstantHDR: Single-forward Gaussian Splatting for High Dynamic Range 3D Reconstruction
por: Ye, Dingqiang, et al.
Publicado: (2026)
por: Ye, Dingqiang, et al.
Publicado: (2026)
Towards Zero-Shot Anomaly Detection and Reasoning with Multimodal Large Language Models
por: Xu, Jiacong, et al.
Publicado: (2025)
por: Xu, Jiacong, et al.
Publicado: (2025)
Diffscaler: Enhancing the Generative Prowess of Diffusion Transformers
por: Nair, Nithin Gopalakrishnan, et al.
Publicado: (2024)
por: Nair, Nithin Gopalakrishnan, et al.
Publicado: (2024)
Morphing Through Time: Diffusion-Based Bridging of Temporal Gaps for Robust Alignment in Change Detection
por: Madani, Seyedehanita, et al.
Publicado: (2025)
por: Madani, Seyedehanita, et al.
Publicado: (2025)
Latent Feature-Guided Diffusion Models for Shadow Removal
por: Mei, Kangfu, et al.
Publicado: (2023)
por: Mei, Kangfu, et al.
Publicado: (2023)
Bigger is not Always Better: Scaling Properties of Latent Diffusion Models
por: Mei, Kangfu, et al.
Publicado: (2024)
por: Mei, Kangfu, et al.
Publicado: (2024)
Scale-Wise VAR is Secretly Discrete Diffusion
por: Kumar, Amandeep, et al.
Publicado: (2025)
por: Kumar, Amandeep, et al.
Publicado: (2025)
Fine-Tuning Image-Conditional Diffusion Models is Easier than You Think
por: Garcia, Gonzalo Martin, et al.
Publicado: (2024)
por: Garcia, Gonzalo Martin, et al.
Publicado: (2024)
Representation Entanglement for Generation: Training Diffusion Transformers Is Much Easier Than You Think
por: Wu, Ge, et al.
Publicado: (2025)
por: Wu, Ge, et al.
Publicado: (2025)
Think Before You Segment: High-Quality Reasoning Segmentation with GPT Chain of Thoughts
por: Kao, Shiu-hong, et al.
Publicado: (2025)
por: Kao, Shiu-hong, et al.
Publicado: (2025)
TPGDiff: Hierarchical Triple-Prior Guided Diffusion for Image Restoration
por: Tu, Yanjie, et al.
Publicado: (2026)
por: Tu, Yanjie, et al.
Publicado: (2026)
DiffRegCD: Integrated Registration and Change Detection with Diffusion Features
por: Madani, Seyedehanita, et al.
Publicado: (2025)
por: Madani, Seyedehanita, et al.
Publicado: (2025)
MaxFusion: Plug&Play Multi-Modal Generation in Text-to-Image Diffusion Models
por: Nair, Nithin Gopalakrishnan, et al.
Publicado: (2024)
por: Nair, Nithin Gopalakrishnan, et al.
Publicado: (2024)
RoboMirror: Understand Before You Imitate for Video to Humanoid Locomotion
por: Li, Zhe, et al.
Publicado: (2025)
por: Li, Zhe, et al.
Publicado: (2025)
Self-Supervised MRI Reconstruction with Unrolled Diffusion Models
por: Korkmaz, Yilmaz, et al.
Publicado: (2023)
por: Korkmaz, Yilmaz, et al.
Publicado: (2023)
From Imitation to Intuition: Intrinsic Reasoning for Open-Instance Video Classification
por: Zhang, Ke, et al.
Publicado: (2026)
por: Zhang, Ke, et al.
Publicado: (2026)
GenDeg: Diffusion-based Degradation Synthesis for Generalizable All-In-One Image Restoration
por: Rajagopalan, Sudarshan, et al.
Publicado: (2024)
por: Rajagopalan, Sudarshan, et al.
Publicado: (2024)
$CrowdDiff$: Multi-hypothesis Crowd Density Estimation using Diffusion Models
por: Ranasinghe, Yasiru, et al.
Publicado: (2023)
por: Ranasinghe, Yasiru, et al.
Publicado: (2023)
MExD: An Expert-Infused Diffusion Model for Whole-Slide Image Classification
por: Zhao, Jianwei, et al.
Publicado: (2025)
por: Zhao, Jianwei, et al.
Publicado: (2025)
DiffDoctor: Diagnosing Image Diffusion Models Before Treating
por: Wang, Yiyang, et al.
Publicado: (2025)
por: Wang, Yiyang, et al.
Publicado: (2025)
Think Before You Drive: World Model-Inspired Multimodal Grounding for Autonomous Vehicles
por: Liao, Haicheng, et al.
Publicado: (2025)
por: Liao, Haicheng, et al.
Publicado: (2025)
Mitigating Long-Tail Bias via Prompt-Controlled Diffusion Augmentation
por: Wijenayake, Buddhi, et al.
Publicado: (2026)
por: Wijenayake, Buddhi, et al.
Publicado: (2026)
Think While You Generate: Discrete Diffusion with Planned Denoising
por: Liu, Sulin, et al.
Publicado: (2024)
por: Liu, Sulin, et al.
Publicado: (2024)
STEREO: A Two-Stage Framework for Adversarially Robust Concept Erasing from Text-to-Image Diffusion Models
por: Srivatsan, Koushik, et al.
Publicado: (2024)
por: Srivatsan, Koushik, et al.
Publicado: (2024)
Ejemplares similares
-
FreeViS: Training-free Video Stylization with Inconsistent References
por: Xu, Jiacong, et al.
Publicado: (2025) -
Endless World: Real-Time 3D-Aware Long Video Generation
por: Zhang, Ke, et al.
Publicado: (2025) -
Wild-GS: Real-Time Novel View Synthesis from Unconstrained Photo Collections
por: Xu, Jiacong, et al.
Publicado: (2024) -
Reference-based Controllable Scene Stylization with Gaussian Splatting
por: Mei, Yiqun, et al.
Publicado: (2024) -
Not All Tokens Need 40 Steps: Heterogeneous Step Allocation in Diffusion Transformers for Efficient Video Generation
por: Chu, Ernie, et al.
Publicado: (2026)