PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff Drop
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Chenyu, Michel, Oscar, Pan, Xichen, Liu, Sainan, Roberts, Mike, Xie, Saining |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Image Sculpting: Precise Object Editing with 3D Geometry Control
by: Yenphraphai, Jiraphon, et al.
Published: (2024)
by: Yenphraphai, Jiraphon, et al.
Published: (2024)
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis
by: Tang, Bingda, et al.
Published: (2025)
by: Tang, Bingda, et al.
Published: (2025)
Cambrian-P: Pose-Grounded Video Understanding
by: Yang, Jihan, et al.
Published: (2026)
by: Yang, Jihan, et al.
Published: (2026)
Repurposing Geometric Foundation Models for Multi-view Diffusion
by: Jang, Wooseok, et al.
Published: (2026)
by: Jang, Wooseok, et al.
Published: (2026)
Solaris: Building a Multiplayer Video World Model in Minecraft
by: Savva, Georgy, et al.
Published: (2026)
by: Savva, Georgy, et al.
Published: (2026)
Deconstructing Denoising Diffusion Models for Self-Supervised Learning
by: Chen, Xinlei, et al.
Published: (2024)
by: Chen, Xinlei, et al.
Published: (2024)
SiT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant Transformers
by: Ma, Nanye, et al.
Published: (2024)
by: Ma, Nanye, et al.
Published: (2024)
Fast Encoding and Decoding for Implicit Video Representation
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
by: Chen, Jiuhai, et al.
Published: (2025)
by: Chen, Jiuhai, et al.
Published: (2025)
Diffusion Transformers with Representation Autoencoders
by: Zheng, Boyang, et al.
Published: (2025)
by: Zheng, Boyang, et al.
Published: (2025)
TPDiff: Temporal Pyramid Video Diffusion Model
by: Ran, Lingmin, et al.
Published: (2025)
by: Ran, Lingmin, et al.
Published: (2025)
Video Diffusion Models are Training-free Motion Interpreter and Controller
by: Xiao, Zeqi, et al.
Published: (2024)
by: Xiao, Zeqi, et al.
Published: (2024)
Morpheus: Benchmarking Physical Reasoning of Video Generative Models with Real Physical Experiments
by: Zhang, Chenyu, et al.
Published: (2025)
by: Zhang, Chenyu, et al.
Published: (2025)
SARD: Segmentation-Aware Anomaly Synthesis via Region-Constrained Diffusion with Discriminative Mask Guidance
by: Wang, Yanshu, et al.
Published: (2025)
by: Wang, Yanshu, et al.
Published: (2025)
Benchmark Designers Should "Train on the Test Set" to Expose Exploitable Non-Visual Shortcuts
by: Brown, Ellis, et al.
Published: (2025)
by: Brown, Ellis, et al.
Published: (2025)
LRQ-DiT: Log-Rotation Post-Training Quantization of Diffusion Transformers for Image and Video Generation
by: Yang, Lianwei, et al.
Published: (2025)
by: Yang, Lianwei, et al.
Published: (2025)
Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
by: Yu, Sihyun, et al.
Published: (2024)
by: Yu, Sihyun, et al.
Published: (2024)
EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos
by: Xu, Jilan, et al.
Published: (2025)
by: Xu, Jilan, et al.
Published: (2025)
DiffusionGuard: A Robust Defense Against Malicious Diffusion-based Image Editing
by: Choi, June Suk, et al.
Published: (2024)
by: Choi, June Suk, et al.
Published: (2024)
MoGAN: Improving Motion Quality in Video Diffusion via Few-Step Motion Adversarial Post-Training
by: Xue, Haotian, et al.
Published: (2025)
by: Xue, Haotian, et al.
Published: (2025)
PISA: Piecewise Sparse Attention Is Wiser for Efficient Diffusion Transformers
by: Li, Haopeng, et al.
Published: (2026)
by: Li, Haopeng, et al.
Published: (2026)
SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding
by: Brown, Ellis, et al.
Published: (2025)
by: Brown, Ellis, et al.
Published: (2025)
SafeWatch: An Efficient Safety-Policy Following Video Guardrail Model with Transparent Explanations
by: Chen, Zhaorun, et al.
Published: (2024)
by: Chen, Zhaorun, et al.
Published: (2024)
Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
by: Tong, Shengbang, et al.
Published: (2024)
by: Tong, Shengbang, et al.
Published: (2024)
QVD: Post-training Quantization for Video Diffusion Models
by: Tian, Shilong, et al.
Published: (2024)
by: Tian, Shilong, et al.
Published: (2024)
Exploring Data-Free LoRA Transferability for Video Diffusion Models
by: Wang, Yuchen, et al.
Published: (2026)
by: Wang, Yuchen, et al.
Published: (2026)
Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models
by: Tang, Yolo Y., et al.
Published: (2025)
by: Tang, Yolo Y., et al.
Published: (2025)
Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
by: Guan, Yiran, et al.
Published: (2026)
by: Guan, Yiran, et al.
Published: (2026)
Scale Where It Matters: Training-Free Localized Scaling for Diffusion Models
by: Ren, Qin, et al.
Published: (2025)
by: Ren, Qin, et al.
Published: (2025)
SeedVR2: One-Step Video Restoration via Diffusion Adversarial Post-Training
by: Wang, Jianyi, et al.
Published: (2025)
by: Wang, Jianyi, et al.
Published: (2025)
Watch Before You Answer: Learning from Visually Grounded Post-Training
by: Zhang, Yuxuan, et al.
Published: (2026)
by: Zhang, Yuxuan, et al.
Published: (2026)
ExVideo: Extending Video Diffusion Models via Parameter-Efficient Post-Tuning
by: Duan, Zhongjie, et al.
Published: (2024)
by: Duan, Zhongjie, et al.
Published: (2024)
Tail-Aware Post-Training Quantization for 3D Geometry Models
by: Pan, Sicheng, et al.
Published: (2026)
by: Pan, Sicheng, et al.
Published: (2026)
Exploring Iterative Refinement with Diffusion Models for Video Grounding
by: Liang, Xiao, et al.
Published: (2023)
by: Liang, Xiao, et al.
Published: (2023)
VipDiff: Towards Coherent and Diverse Video Inpainting via Training-free Denoising Diffusion Models
by: Xie, Chaohao, et al.
Published: (2025)
by: Xie, Chaohao, et al.
Published: (2025)
Magic Fixup: Streamlining Photo Editing by Watching Dynamic Videos
by: Alzayer, Hadi, et al.
Published: (2024)
by: Alzayer, Hadi, et al.
Published: (2024)
WAT: Online Video Understanding Needs Watching Before Thinking
by: Han, Zifan, et al.
Published: (2026)
by: Han, Zifan, et al.
Published: (2026)
Self-Refining Video Sampling
by: Jang, Sangwon, et al.
Published: (2026)
by: Jang, Sangwon, et al.
Published: (2026)
Learning A Physical-aware Diffusion Model Based on Transformer for Underwater Image Enhancement
by: Zhao, Chen, et al.
Published: (2024)
by: Zhao, Chen, et al.
Published: (2024)
Transfer between Modalities with MetaQueries
by: Pan, Xichen, et al.
Published: (2025)
by: Pan, Xichen, et al.
Published: (2025)
Similar Items
-
Image Sculpting: Precise Object Editing with 3D Geometry Control
by: Yenphraphai, Jiraphon, et al.
Published: (2024) -
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis
by: Tang, Bingda, et al.
Published: (2025) -
Cambrian-P: Pose-Grounded Video Understanding
by: Yang, Jihan, et al.
Published: (2026) -
Repurposing Geometric Foundation Models for Multi-view Diffusion
by: Jang, Wooseok, et al.
Published: (2026) -
Solaris: Building a Multiplayer Video World Model in Minecraft
by: Savva, Georgy, et al.
Published: (2026)